In plain words: A textbook covering the core ideas behind large language models: how they are pre-trained, generate text, follow prompts, get steered toward human preferences, and reason. It teaches foundations rather than the newest techniques, serving as a reference for students and practitioners.
Abstract
This is a book about large language models. As indicated by the title, it primarily focuses on foundational concepts rather than comprehensive coverage of all cutting-edge technologies. The book is structured into six main chapters, each exploring a key area: pre-training, generative models, prompting, alignment, inference, and reasoning. It is intended for college students, professionals, and practitioners in natural language processing and related fields, and can serve as a reference for anyone interested in large language models.
Tong Xiao, Jingbo Zhu
arXiv:2501.09223 · cs.CL, cs.AI, cs.LG · submitted Jan 16, 2025 · updated Sep 24, 2026
abstract · pdf · html · Added a new chapter
Assume you are a college instructor for a Freshman Computer Science course.
Your job is to take a pdf file from the internet and teach the topics to you students.
You will do this by writing paragraphs or bullet points about any and all key concepts in the PDF necessary to cover the topic in 2 hours of lectures
The pdf file is at https://arxiv.org/pdf/2501.09223
Build the lecture for me.