In plain words: A textbook covering the core ideas behind large language models: how they are pre-trained, generate text, follow prompts, get steered toward human preferences, and reason. It teaches foundations rather than the newest techniques, serving as a reference for students and practitioners.
Abstract
This is a book about large language models. As indicated by the title, it primarily focuses on foundational concepts rather than comprehensive coverage of all cutting-edge technologies. The book is structured into six main chapters, each exploring a key area: pre-training, generative models, prompting, alignment, inference, and reasoning. It is intended for college students, professionals, and practitioners in natural language processing and related fields, and can serve as a reference for anyone interested in large language models.
Tong Xiao, Jingbo Zhu
arXiv:2501.09223 · cs.CL, cs.AI, cs.LG · submitted Jan 16, 2025 · updated Sep 24, 2026
abstract · pdf · html · Added a new chapter