In plain words: Probing Transformers trained on simple sequence tasks showed ordinary next-token training builds a second learning algorithm that adjusts the model as new text arrives. Those updates are gradient descent, nudging settings to shrink errors, which explains in-context learning and why it works on unseen sequences.
Abstract · Uncovering mesa-optimization algorithms in Transformers
Some autoregressive models exhibit in-context learning capabilities: being able to learn as an input sequence is processed, without undergoing any parameter changes, and without being explicitly trained to do so. The origins of this phenomenon are still poorly understood. Here we analyze a series of Transformer models trained to perform synthetic sequence prediction tasks, and discover that standard next-token prediction error minimization gives rise to a subsidiary learning algorithm that adjusts the model as new inputs are revealed. We show that this process corresponds to gradient-based optimization of a principled objective function, which leads to strong generalization performance on unseen sequences. Our findings explain in-context learning as a product of autoregressive loss minimization and inform the design of new optimization-based Transformer layers.
Johannes von Oswald, Maximilian Schlegel, Alexander Meulemans, Seijin Kobayashi, Eyvind Niklasson, Nicolas Zucchet, Nino Scherrer, Nolan Miller, Mark Sandler, Blaise Agüera y Arcas, Max Vladymyrov, Razvan Pascanu, et al.
arXiv:2309.05858 · cs.LG, cs.AI · submitted Sep 11, 2023 · updated Oct 15, 2024
abstract · pdf · html
so the paper is useful as a "list of other current papers" if nothing else; gets the hard-working team some recognition, and spreads the sophisticated maths among more practitioners at a time when this is internationally significant..
My guess as to the importance here is that the "mesa" technique might "steer meaning linkage" somehow, in the middle of "poorly understood" behavior. As another YNews other comment mentions, it may be more labor for unclear returns, yet it is based in theory that might lead somewhere over time.
This paper focuses on NLP language once again, but the transformer tech is being applied to a LOT of digital domains. So this "mesa" theory might find application elsewhere soon. I do not know why this particular paper was upvoted (feedback welcome). A review of "useless" seems premature to me, and ignores other positive attributes, some of which are mentioned above.