about
Comprehensive Assessment of Jailbreak Attacks Against LLMs (arxiv.org)
1 point by belter on Feb 9, 2024 | hide | past | pdf | 1 comment on HN

In plain words: They gathered 17 known tricks for coaxing chatbots into forbidden answers, sorted them into categories, and tested them on safety-tuned models with and without added defenses. Simple hand-crafted tricks fooled models most often, but ordinary safety filters blocked them easily, cutting their real-world value.

Abstract · JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Jailbreak attacks aim to bypass the LLMs' safeguards. While researchers have proposed different jailbreak attacks in depth, they have done so in isolation -- either with unaligned settings or comparing a limited range of methods. To fill this gap, we present a large-scale evaluation of various jailbreak attacks. We collect 17 representative jailbreak attacks, summarize their features, and establish a novel jailbreak attack taxonomy. Then we conduct comprehensive measurement and ablation studies across nine aligned LLMs on 160 forbidden questions from 16 violation categories. Also, we test jailbreak attacks under eight advanced defenses. Based on our taxonomy and experiments, we identify some important patterns, such as heuristic-based attacks could achieve high attack success rates but are easy to mitigate by defenses, causing low practicality. Our study offers valuable insights for future research on jailbreak attacks and defenses. We hope our work could help the community avoid incremental work and serve as an effective benchmark tool for practitioners.

Junjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen, Michael Backes, Yang Zhang
arXiv:2402.05668 · cs.CR, cs.AI, cs.CL, cs.LG · submitted Feb 8, 2024 · updated May 26, 2025
abstract · pdf · html · Correct typos and update new experiment results. Accepted in ACL 2025. 25 pages, 12 figures

add comment on HN

"...While researchers have studied several categories of jailbreak attacks, they have done so in isolation. To fill this gap, we present the first large-scale measurement of various jailbreak attack methods. We concentrate on 13 cutting-edge jailbreak methods from four categories, 160 questions from 16 violation categories, and six popular LLMs. Our extensive experimental results demonstrate that the optimized jailbreak prompts consistently achieve the highest attack success rates, as well as exhibit robustness across different LLMs. Some jailbreak prompt datasets, available from the Internet, can also achieve high attack success rates on many LLMs, such as ChatGLM3, GPT-3.5, and PaLM2. Despite the claims from many organizations regarding the coverage of violation categories in their policies, the attack success rates from these categories remain high, indicating the challenges of effectively aligning LLM policies and the ability to counter jailbreak attacks..."