In plain words: A wide-ranging survey of reinforcement learning, where an agent learns through trial, error, and rewards to make a sequence of decisions. It walks through the main families of techniques, from classic control methods to using rewards to train large language models.
Abstract
This manuscript gives a big-picture, up-to-date overview of the field of (deep) reinforcement learning and sequential decision making, covering value-based methods, policy-based methods, model-based methods, multi-agent RL, LLMs and RL, and various other topics (e.g., offline RL, hierarchical RL, intrinsic reward). It also includes some code snippets for training LLMs with RL.
Kevin Murphy
arXiv:2412.05265 · cs.AI, cs.LG · submitted Dec 6, 2024 · updated Dec 1, 2025
abstract · pdf · html
I struggled a lot with the first chapter, and had to look up a lot of terms that weren’t defined. But ultimately it was one of the most worthwhile things I read, and has helped me follow along with other important papers.