In plain words: Course notes from an ETH Zürich class lay out the math behind large language models: what formally counts as a language model and how one is built. They cover the theory side of the course, offering foundations rather than a new experiment or result.
Abstract
Large language models have become one of the most commonly deployed NLP inventions. In the past half-decade, their integration into core natural language processing tools has dramatically increased the performance of such tools, and they have entered the public discourse surrounding artificial intelligence. Consequently, it is important for both developers and researchers alike to understand the mathematical foundations of large language models, as well as how to implement them. These notes are the accompaniment to the theoretical portion of the ETH Zürich course on large language models, covering what constitutes a language model from a formal, theoretical perspective.
Ryan Cotterell, Anej Svete, Clara Meister, Tianyu Liu, Li Du
arXiv:2311.04329 · cs.CL · submitted Nov 7, 2023 · updated Apr 17, 2024
abstract · pdf · html