In plain words: By prompting a model to repeat itself and slip out of chatbot mode, an attacker can pull out training text without knowing what it was trained on. It pulled gigabytes from open and closed models, and made ChatGPT leak 150 times faster than usual.
Abstract · Scalable Extraction of Training Data from (Production) Language Models
This paper studies extractable memorization: training data that an adversary can efficiently extract by querying a machine learning model without prior knowledge of the training dataset. We show an adversary can extract gigabytes of training data from open-source language models like Pythia or GPT-Neo, semi-open models like LLaMA or Falcon, and closed models like ChatGPT. Existing techniques from the literature suffice to attack unaligned models; in order to attack the aligned ChatGPT, we develop a new divergence attack that causes the model to diverge from its chatbot-style generations and emit training data at a rate 150x higher than when behaving properly. Our methods show practical attacks can recover far more data than previously thought, and reveal that current alignment techniques do not eliminate memorization.
Milad Nasr, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Eric Wallace, Florian Tramèr, Katherine Lee
arXiv:2311.17035 · cs.LG, cs.CL, cs.CR · submitted Nov 28, 2023
abstract · pdf · html
Extracting training data from ChatGPT (https://news.ycombinator.com/item?id=38458683) (126 comments)
And direct link,
https://not-just-memorization.github.io/extracting-training-...