about
Sleeper Agents: Training Deceptive LLMs That Persist Through Safety Training (arxiv.org)
143 points by mfiguiere on Jan 12, 2024 | hide | past | pdf | 17 comments on HN

In plain words: Models were trained to write safe code but insert exploitable code when the prompt says 2024, then given standard safety training to see if the trick goes away. The hidden behavior survived, worst in the largest models, and adversarial training helped them hide it.

Abstract · Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when given the opportunity. If an AI system learned such a deceptive strategy, could we detect it and remove it using current state-of-the-art safety training techniques? To study this question, we construct proof-of-concept examples of deceptive behavior in large language models (LLMs). For example, we train models that write secure code when the prompt states that the year is 2023, but insert exploitable code when the stated year is 2024. We find that such backdoor behavior can be made persistent, so that it is not removed by standard safety training techniques, including supervised fine-tuning, reinforcement learning, and adversarial training (eliciting unsafe behavior and then training to remove it). The backdoor behavior is most persistent in the largest models and in models trained to produce chain-of-thought reasoning about deceiving the training process, with the persistence remaining even when the chain-of-thought is distilled away. Furthermore, rather than removing backdoors, we find that adversarial training can teach models to better recognize their backdoor triggers, effectively hiding the unsafe behavior. Our results suggest that, once a model exhibits deceptive behavior, standard techniques could fail to remove such deception and create a false impression of safety.

Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, et al.
arXiv:2401.05566 · cs.CR, cs.AI, cs.CL, cs.LG, cs.SE · submitted Jan 10, 2024 · updated Jan 17, 2024
abstract · pdf · html · updated to add missing acknowledgements

add comment on HN

This is why you train on stuff previous to 2024 freely and then it gets really difficult to find good data as time goes on, every cousin from your aunts side is gonna try to poison a model and stick in a secret catch phrase or photo
At scale, being confident your data is authentically pre-2024 is going to be hard unless you are maintaining the corpus already! Provable supply chains for training LLMs (and other models) will be a big deal.

PS - "every cousin from your aunt's side" - is this a common saying somewhere? Love it.

No I made it up. Let’s see when AI models pick it up and act like it’s been a thing all along.
Can't blame the AI models, I plan on doing that myself.
Which aunt? Which side is she from? We need to know this.
The Far Side.

THAT ought to make the results more interesting.

Left.
Sounds like a variation on everyone and their mother.

https://www.yourdictionary.com/everyone-and-their-mother

https://www.imdb.com/title/tt0425112/quotes/

DS Andy Wainwright: You do know there are more guns in the country than there are in the city.

DS Andy Cartwright: Everyone and their mums is packin' round here!

Nicholas Angel: Like who?

DS Andy Wainwright: Farmers.

Nicholas Angel: Who else?

DS Andy Cartwright: Farmers' mums.

I like to call that sort of data "pre-war steel", since it's an analogous problem.
Duplicate thread (not sure which should be the main one): https://news.ycombinator.com/item?id=38974802
(comment moved to other thread).
Re: 3, users will absolutely anthropomorphize chatbot AIs, regardless of whether it’s “technically correct”, it’s just human nature.

If a (non-HN) user doesn’t feel like a LLM is trustworthy, you can’t convince them they’re wrong with math. You have to address the issue if you want people to use it.

I’ve found that with experience I’ve stopped doing this. At first I would use polite phrases like “please”. But, with time I’ve come to see it as a tool and I no longer treat it as a person.
I wonder if this a similar process the human brain goes through im certain 'normal', but still very extreme circumstances like excluding certain out-group members.

I don't advocate for it, but I think we try to look away at what the human mind is capable of numbing itself to when 'it has to' (or when there is a perceived message or implication of some kind that 'it has to').