In plain words: Because language exists to convey meaning, the word–meaning pairs it covers are rare and strongly peaked, and huge language models learn exactly those peaks. Understanding, in-context learning, chain-of-thought, and instruction tuning all follow from one rule: weighing possible meanings by how likely they are.
Abstract
Languages are not created randomly but rather to communicate information. There is a strong association between languages and their underlying meanings, resulting in a sparse joint distribution that is heavily peaked according to their correlations. Moreover, these peak values happen to match with the marginal distribution of languages due to the sparsity. With the advent of LLMs trained on big data and large models, we can now precisely assess the marginal distribution of languages, providing a convenient means of exploring the sparse structures in the joint distribution for effective inferences. In this paper, we categorize languages as either unambiguous or ε-ambiguous and present quantitative results to demonstrate that the emergent abilities of LLMs, such as language understanding, in-context learning, chain-of-thought prompting, and effective instruction fine-tuning, can all be attributed to Bayesian inference on the sparse joint distribution of languages.
Hui Jiang
arXiv:2304.09960 · cs.CL, cs.AI, cs.LG · submitted Apr 19, 2023 · updated Sep 13, 2023
abstract · pdf · html · 17 pages, 3 figures
> Languages are not created randomly, but with a specific purpose in mind, which is to convey information. Languages are composed of distinct, relatively independent units, such as sentences in natural languages or statements in programming languages. These separate pieces of language are referred to as "messages" in this paper. Each message, represented as x, is in turn composed of a sequence of symbols from an alphabet with varying lengths. A message is created with the aim of expressing a single and definite intention, denoted as θ. The set of all possible intentions constitutes another space, denoted as Θ. We assume that the intention space Θ is a countable set of many distinct intentions. Each θ may represent a simple intention, which is an element from a finite set, or a composite intention that is made up of several simpler concepts or components through concatenation or recursion. Here we only require that the intention space Θ is discrete and complete, and each element in Θ is unique.