about
Advancing Language Model Reasoning Through RL and Inference Scaling (arxiv.org)
2 points by frozenseven on Jan 25, 2025 | hide | past | pdf | discuss on HN

In plain words: T1 trains a language model to solve hard math problems by first practicing trial-and-error answers, then learning from rewards while trying many different solutions. On tough math benchmarks, letting it think longer kept boosting accuracy without a separate checker, unlike usual imitation-trained models.

Abstract · T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Large language models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks. However, existing approaches mainly rely on imitation learning and struggle to achieve effective test-time scaling. While reinforcement learning (RL) holds promise for enabling self-exploration, recent attempts yield modest improvements in complex reasoning. In this paper, we present T1 to scale RL by encouraging exploration and understand inference scaling. We first initialize the LLM using synthesized chain-of-thought data that integrates trial-and-error and self-verification. To scale RL training, we promote increased sampling diversity through oversampling. We demonstrate that T1 with open LLMs as its base exhibits inference scaling behavior and achieves superior performance on challenging math reasoning benchmarks. More importantly, we present a simple strategy to examine inference scaling, where increased inference budgets directly lead to T1's better performance without any additional verification.

Zhenyu Hou, Xin Lv, Rui Lu, Jiajie Zhang, Yujiang Li, Zijun Yao, Juanzi Li, Jie Tang, Yuxiao Dong
arXiv:2501.11651 · cs.LG, cs.CL · submitted Jan 20, 2025 · updated Jun 13, 2025
abstract · pdf · html · Accepted to ICML 2025

add comment on HN