In plain words: A new model treats a text generator as a big table of next-word chances, with a starting guess updated as text arrives. It shows large models learn this way, explaining why bigger ones pick up patterns from prompt examples, matching measured next-word probabilities.
Abstract · Beyond the Black Box: A Statistical Model for LLM Reasoning and Inference
This paper introduces a novel Bayesian learning model to explain the behavior of Large Language Models (LLMs), focusing on their core optimization metric of next token prediction. We develop a theoretical framework based on an ideal generative text model represented by a multinomial transition probability matrix with a prior, and examine how LLMs approximate this matrix. Key contributions include: (i) a continuity theorem relating embeddings to multinomial distributions, (ii) a demonstration that LLM text generation aligns with Bayesian learning principles, (iii) an explanation for the emergence of in-context learning in larger models, (iv) empirical validation using visualizations of next token probabilities from an instrumented Llama model Our findings provide new insights into LLM functioning, offering a statistical foundation for understanding their capabilities and limitations. This framework has implications for LLM design, training, and application, potentially guiding future developments in the field.
Siddhartha Dalal, Vishal Misra
arXiv:2402.03175 · cs.LG, cs.AI · submitted Feb 5, 2024 · updated Sep 24, 2024
abstract · pdf · html
In this paper we present a new model to explain the behavior of Large Language Models. Our frame of reference is an abstract probability matrix, which contains the multinomial probabilities for next token prediction in each row, where the row represents a specific prompt. We then demonstrate that LLM text generation is consistent with a compact representation of this abstract matrix through a combination of embeddings and Bayesian learning. Our model explains (the emergence of) In-Context learning with scale of the LLMs, as also other phenomena like Chain of Thought reasoning and the problem with large context windows. Finally, we outline implications of our model and some directions for future exploration.
Where does the "Cannot Recursively Improve" come from?