about
Defeating the Training-Inference Mismatch via FP16 (arxiv.org)
1 point by billyzs 335 days ago | hide | past | pdf | discuss on HN

In plain words: When a language model is fine-tuned with rewards, the common BF16 number format rounds values so coarsely that it answers differently in training and use, making learning unstable. Switching to the more precise FP16 format closes that gap, giving steadier, faster, stronger results.

Abstract

Reinforcement learning (RL) fine-tuning of large language models (LLMs) often suffers from instability due to the numerical mismatch between the training and inference policies. While prior work has attempted to mitigate this issue through algorithmic corrections or engineering alignments, we show that its root cause lies in the floating point precision itself. The widely adopted BF16, despite its large dynamic range, introduces large rounding errors that breaks the consistency between training and inference. In this work, we demonstrate that simply reverting to \textbf{FP16} effectively eliminates this mismatch. The change is simple, fully supported by modern frameworks with only a few lines of code change, and requires no modification to the model architecture or learning algorithm. Our results suggest that using FP16 uniformly yields more stable optimization, faster convergence, and stronger performance across diverse tasks, algorithms and frameworks. We hope these findings motivate a broader reconsideration of precision trade-offs in RL fine-tuning.

Penghui Qi, Zichen Liu, Xiangxin Zhou, Tianyu Pang, Chao Du, Wee Sun Lee, Min Lin
arXiv:2510.26788 · cs.LG, cs.AI, cs.CL · submitted Oct 30, 2025
abstract · pdf · html

add comment on HN
Also discussed: Oct 2025 (2 points, 0 comments)