about
Large Language Models (arxiv.org)
3 points by belter on Jul 18, 2023 | hide | past | pdf | 1 comment on HN

In plain words: These lectures trace how large language models developed and explain the transformer design behind them, for readers with math or physics backgrounds. They survey current ideas on why predicting the next word lets a model handle many other intelligent tasks.

Abstract

Artificial intelligence is making spectacular progress, and one of the best examples is the development of large language models (LLMs) such as OpenAI's GPT series. In these lectures, written for readers with a background in mathematics or physics, we give a brief history and survey of the state of the art, and describe the underlying transformer architecture in detail. We then explore some current ideas on how LLMs work and how models trained to predict the next word in a text are able to perform other tasks displaying intelligence.

Michael R. Douglas
arXiv:2307.05782 · cs.CL, hep-th, math.HO, physics.comp-ph · submitted Jul 11, 2023 · updated Oct 6, 2023
abstract · pdf · html · 47 pages (v2: added references, corrected typos)

add comment on HN
Also discussed: Oct 2023 (1 point, 0 comments)

"...In these lectures, written for readers with a background in mathematics or physics, we give a brief history and survey of the state of the art, and describe the underlying transformer architecture in detail. We then explore some current ideas on how LLMs work and how models trained to predict the next word in a text are able to perform other tasks displaying intelligence..."