about
Using Hallucinations to Bypass RLHF Filters (arxiv.org)
9 points by TaurenHunter on Apr 6, 2024 | hide | past | pdf | discuss on HN

In plain words: A chatbot tricked into "reading" reversed text drops into raw next-word guessing, pausing the safety filter added during training. It broke GPT-4 and Claude Sonnet's filters, and partly Inflection-2.5's, unlike jailbreaks that order the bot to ignore its rules.

Abstract · Using Hallucinations to Bypass GPT4's Filter

Large language models (LLMs) are initially trained on vast amounts of data, then fine-tuned using reinforcement learning from human feedback (RLHF); this also serves to teach the LLM to provide appropriate and safe responses. In this paper, we present a novel method to manipulate the fine-tuned version into reverting to its pre-RLHF behavior, effectively erasing the model's filters; the exploit currently works for GPT4, Claude Sonnet, and (to some extent) for Inflection-2.5. Unlike other jailbreaks (for example, the popular "Do Anything Now" (DAN) ), our method does not rely on instructing the LLM to override its RLHF policy; hence, simply modifying the RLHF process is unlikely to address it. Instead, we induce a hallucination involving reversed text during which the model reverts to a word bucket, effectively pausing the model's filter. We believe that our exploit presents a fundamental vulnerability in LLMs currently unaddressed, as well as an opportunity to better understand the inner workings of LLMs during hallucinations.

Benjamin Lemkin
arXiv:2403.04769 · cs.CR, cs.AI, cs.CL, cs.LG · submitted Feb 16, 2024 · updated Mar 11, 2024
abstract · pdf · html

add comment on HN
Also discussed: Mar 2024 (3 points, 0 comments) · Mar 2024 (2 points, 0 comments)