about
A Primer in Post-Training Reasoning Data: What We Know About How It Works (arxiv.org)
2 points by Anon84 122 days ago | hide | past | pdf | discuss on HN

In plain words: It gathers over 150 studies on data used to teach reasoning models after initial training, sorting it by what it is, what makes it useful, how it is built, and how it scales. It turns scattered findings into one framework for judging data releases.

Abstract

Post-training has become a primary driver of recent progress in large reasoning models, and reasoning data are often the key variable determining whether this stage succeeds. Work on post-training reasoning data has grown rapidly, yet this literature remains scattered across dataset papers, reinforcement-learning recipes, reward-model studies, benchmarks, and frontier system reports. This paper is the first primer to synthesize over 150 key public studies and system reports on post-training reasoning data. We organize the field around four questions: what data objects exist, what makes them useful, how they are constructed, and how they scale. Together, this organization provides an attribution framework for future reasoning-data releases and post-training recipes.

Yaoming Li, Guangxiang Zhao, Qilong Shi, Lin Sun, Xiangzheng Zhang, Tong Yang
arXiv:2606.02113 · cs.CL, cs.AI · submitted Jun 1, 2026
abstract · pdf · html · 22 pages. Project Repository: https://github.com/RenBing-Sumeru/Awesome-LLM-Reasoning-Data

add comment on HN