In plain words: A theory shows that even a simple linear model trained to predict the next word in step-by-step reasoning can imitate any computation a computer can do efficiently. Experiments confirm tiny networks handle text and arithmetic well, suggesting the training scheme matters more than the architecture.
Abstract · Auto-Regressive Next-Token Predictors are Universal Learners
Large language models display remarkable capabilities in logical and mathematical reasoning, allowing them to solve complex tasks. Interestingly, these abilities emerge in networks trained on the simple task of next-token prediction. In this work, we present a theoretical framework for studying auto-regressive next-token predictors. We demonstrate that even simple models such as linear next-token predictors, trained on Chain-of-Thought (CoT) data, can approximate any function efficiently computed by a Turing machine. We introduce a new complexity measure -- length complexity -- which measures the number of intermediate tokens in a CoT sequence required to approximate some target function, and analyze the interplay between length complexity and other notions of complexity. Finally, we show experimentally that simple next-token predictors, such as linear networks and shallow Multi-Layer Perceptrons (MLPs), display non-trivial performance on text generation and arithmetic tasks. Our results demonstrate that the power of today's LLMs can be attributed, to a great extent, to the auto-regressive next-token training scheme, and not necessarily to a particular choice of architecture.
Eran Malach
arXiv:2309.06979 · cs.LG, cs.CL · submitted Sep 13, 2023 · updated Jul 29, 2024
abstract · pdf · html
I would have hoped they would attribute LLM success to the structure of language itself. As the authors say, even small linear models can approximate CoT and solve complex tasks. So it's not the model. It's the data.
Analogously, humans have very different brains when looking at low level, but brains still learn the same language and skills about as good as any other brain. It's not the brain or neural net (the models) but the data that shapes them to become smart.
This insight has consequences on how we view training data and where to focus our work to improve AI and human brains - improve language, ideas & chains of thought. This resonates with recent discoveries in fine-tuning and training small models like phi-1 and phi-1.5 who were trained on "textbook quality" data of high diversity.