about
Rethinking Reflection in Pre-Training (arxiv.org)
1 point by swyx on Apr 8, 2025 | hide | past | pdf | discuss on HN

In plain words: They slipped deliberate mistakes into written-out reasoning and checked whether the model could spot and fix them to reach the right answer. This self-correcting skill showed up early in pre-training and kept improving, rather than appearing only later during reinforcement learning.

Abstract

A language model's ability to reflect on its own reasoning provides a key advantage for solving complex problems. While most recent research has focused on how this ability develops during reinforcement learning, we show that it actually begins to emerge much earlier - during the model's pre-training. To study this, we introduce deliberate errors into chains-of-thought and test whether the model can still arrive at the correct answer by recognizing and correcting these mistakes. By tracking performance across different stages of pre-training, we observe that this self-correcting ability appears early and improves steadily over time. For instance, an OLMo2-7B model pre-trained on 4 trillion tokens displays self-correction on our six self-reflection tasks.

Essential AI, :, Darsh J Shah, Peter Rushton, Somanshu Singla, Mohit Parmar, Kurt Smith, Yash Vanjani, Ashish Vaswani, Adarsh Chaluvaraju, Andrew Hojel, Andrew Ma, et al.
arXiv:2504.04022 · cs.CL, cs.AI · submitted Apr 5, 2025
abstract · pdf · html

add comment on HN