about
Foundations of Large Language Models (arxiv.org)
219 points by pkoird on Jan 23, 2025 | hide | past | pdf | 20 comments on HN

In plain words: A textbook covering the core ideas behind large language models: how they are pre-trained, generate text, follow prompts, get steered toward human preferences, and reason. It teaches foundations rather than the newest techniques, serving as a reference for students and practitioners.

Abstract

This is a book about large language models. As indicated by the title, it primarily focuses on foundational concepts rather than comprehensive coverage of all cutting-edge technologies. The book is structured into six main chapters, each exploring a key area: pre-training, generative models, prompting, alignment, inference, and reasoning. It is intended for college students, professionals, and practitioners in natural language processing and related fields, and can serve as a reference for anyone interested in large language models.

Tong Xiao, Jingbo Zhu
arXiv:2501.09223 · cs.CL, cs.AI, cs.LG · submitted Jan 16, 2025 · updated Sep 24, 2026
abstract · pdf · html · Added a new chapter

add comment on HN
Also discussed: Sep 2026 (3 points, 1 comment) · Jan 2025 (3 points, 0 comments)

Me to ChatGPT:

Assume you are a college instructor for a Freshman Computer Science course.

Your job is to take a pdf file from the internet and teach the topics to you students.

You will do this by writing paragraphs or bullet points about any and all key concepts in the PDF necessary to cover the topic in 2 hours of lectures

The pdf file is at https://arxiv.org/pdf/2501.09223

Build the lecture for me.

It can't read PDFs. If you ask it to, it generates code to read the first X characters of the PDF and does a bad job.

(Claude is much better at it.)

Yes it can - both via websearches and uploaded (atleast I'm doing it daily).

EDIT: This article says its only in ChatGPT Enterprise, but works for me on free plan: https://help.openai.com/en/articles/10416312-visual-retrieva...

That article is referencing visuals embedded in PDFs. As a free user you wouldn't be able to ask ChatGPT to analyze a graph inside a PDF, only text.
Authors are from Northeaster University, Shenyang, China, not the Northeastern U in Boston. Don't understand why the two Chinese professors write an LLM book in english, definitely not from experiences, probably under pressure to publish.
prob. not prof; just phd students needs pubs to graduate
These things can be on Arxiv??
I assumed Arxiv was peer-reviewed content only, but it looks like that is not the case.

Submission guidelines: https://info.arxiv.org/help/submit/index.html

Moderation process: https://info.arxiv.org/help/moderation/index.html

On the contrary, ArXiv is for pre-prints, i.e. not (yet) peer-reviewed. Off the top of the my head, it was initially used by physicists who often have huge collaborations and long reviewing time. Then the ML community invaded the space later on. This does not mean a peer-reviewed paper cannot go there of course.
Most of the times the peer review version has a copiright restriction, so the arxiv version is the finañ draft that may have small differences.
As an academic, I always thought of arxiv as where you put your papers first, before they are peer reviewed. Before that we used our webpages, but they kept breaking.
Didn't know I could find it on arxiv, will definitely give it a read
at 231 pages this is definitely book territory
Thankfully the submission is self aware- the first sentence of the article is literally:

> This is a book about large language models.

The book too it self aware, though you do have to make it to page ii.

> In writing this book, we have gradually realized that it is more like a compilation of "notes" we have taken while learning about large language models. Through this note-taking writing style, we hope to offer readers a flexible learning path. Whether they wish to dive deep into a specific area or gain a comprehensive understanding of large language models, they will find the knowledge and insights they need within these "notes".

Now I wonder if the LLMs described in it are self-aware too, and whether by the time I reach the end of this book, I will become self-aware as well.
Is it just me or this book looks rather like a Word doc than a Latex one?
A- Who cares?

B- The latex source of the book is available on the ArXiv page.

It's just you