In plain words: A mathematically precise guide to how transformer models — the design behind today's language models — are built, trained, and used, covering their key parts and best-known versions. It gives exact definitions and equations instead of new results or experiments.
Abstract
This document aims to be a self-contained, mathematically precise overview of transformer architectures and algorithms (*not* results). It covers what transformers are, how they are trained, what they are used for, their key architectural components, and a preview of the most prominent models. The reader is assumed to be familiar with basic ML terminology and simpler neural network architectures such as MLPs.
Mary Phuong, Marcus Hutter
arXiv:2207.09238 · cs.LG, cs.AI, cs.CL, cs.NE · submitted Jul 19, 2022
abstract · pdf · html · 16 pages, 15 algorithms