about
SFT matches RL if you MCMC the training data first (arxiv.org)
2 points by mrkn1 1 day ago | hide | past | pdf | discuss on HN

In plain words: Instead of changing the learning rule, they reshape expert example data with a sampling trick that gradually makes it look like the model's own output. Plain supervised training on this data then matches reinforcement-style training, often generalizing better and forgetting less.

Abstract · Finetuning with Sampling: SFT Learns Better Than You Think

Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.

Aayush Karan, Sitan Chen, Yilun Du
arXiv:2610.02140 · cs.LG, cs.AI, cs.CL · submitted Oct 1, 2026
abstract · pdf · html

add comment on HN