In plain words: A language model was repeatedly trained on text it wrote itself, like a student studying only their own past answers. After several rounds its writing got worse and collapsed into repetitive words.
Abstract · Collapse of Self-trained Language Models
In various fields of knowledge creation, including science, new ideas often build on pre-existing information. In this work, we explore this concept within the context of language models. Specifically, we explore the potential of self-training models on their own outputs, akin to how humans learn and build on their previous thoughts and actions. While this approach is intuitively appealing, our research reveals its practical limitations. We find that extended self-training of the GPT-2 model leads to a significant degradation in performance, resulting in repetitive and collapsed token output.
David Herel, Tomas Mikolov
arXiv:2404.02305 · cs.CL, cs.AI · submitted Apr 2, 2024
abstract · pdf · html · ICLR 2024
The reason why unanchored training fails is fairly simple. "Training" is a misnomer, we're really copying and compressing[1]. When you train a model on itself, you're making a lossy copy of the original, which isn't a very good truth anchor.
There's probably other ways to anchor a self-training process, though. ChatGPT and other text-to-text transformer models are operated as autoregressive processes, where the model spits out a probability distribution, which you then sample to get a token to add to the input, and then repeat until the model says stop. You'll notice that if you squint a little, this looks like the policy function of AlphaGo, but being run stochastically instead of being min-maxed. Which begs the question: why can't we train GPT like we train chess AI, with self-play followed by fine-tuning on the result, as scored by some kind of reward model?
Granted, you'd have to specify a reward model, as well as what behavior you're trying to 'reward'. One other idea that's been bouncing around my head for self-training is training a model to remember details of prior conversations that have since fell off the end of the context window. The biological analogy being "long-term memory", in contrast to the "short-term memory" of the context window. So perhaps your reward model is the model plus the current context window, and your loss is calculated on the same model but without the parts of the context window you want to free up.
No clue if this has already been done, but if it has please reply with the name of the thing I'm not aware of.
[0] Or difference, I forget. If you get the signs wrong you get a hilariously horny version of ChatGPT.
[1] And, thanks to induction heads, compressing the knowledge of how and what to copy.