about
Paper on turning single-shot jailbreaks to multi-shot jailbreaks (arxiv.org)
4 points by IEatPrompts on Sep 3, 2024 | hide | past | pdf | discuss on HN

In plain words: Instead of asking a chatbot one obviously harmful question, this tool splits it into harmless-looking sub-questions asked over several turns until the answer is assembled. It raised how often four AI chatbots complied by up to 46.22% over standard attacks.

Abstract · FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks)

This paper introduces FRACTURED-SORRY-Bench, a framework for evaluating the safety of Large Language Models (LLMs) against multi-turn conversational attacks. Building upon the SORRY-Bench dataset, we propose a simple yet effective method for generating adversarial prompts by breaking down harmful queries into seemingly innocuous sub-questions. Our approach achieves a maximum increase of +46.22\% in Attack Success Rates (ASRs) across GPT-4, GPT-4o, GPT-4o-mini, and GPT-3.5-Turbo models compared to baseline methods. We demonstrate that this technique poses a challenge to current LLM safety measures and highlights the need for more robust defenses against subtle, multi-turn attacks.

Aman Priyanshu, Supriti Vijay
arXiv:2408.16163 · cs.CL, cs.AI · submitted Aug 28, 2024 · updated Nov 7, 2024
abstract · pdf · html · 4 pages, 2 tables

add comment on HN