In plain words: Language models forecast real events and are graded only once those events happen, so the future supplies free answers instead of human labelers. That cut forecast error by 27% and halved how often its stated odds were off, beating a far larger model.
Abstract
Time creates free supervision: forecasts about real-world events resolve to verifiable outcomes. The passage of time provides labels that require no annotation. To exploit this structure, we extend reinforcement learning with verifiable rewards to real-world prediction over time. We train language models to make probabilistic forecasts from causally masked information, using proper scoring rules as the reward function once events resolve. Learning is driven entirely by realized outcomes, enabling scalable outcome-based supervision in open-world prediction. On real-world forecasting benchmarks, Qwen3-32B trained using Foresight Learning improves Brier score by 27% and halves calibration error relative to its pretrained baseline, and outperforms Qwen3-235B on both constructed future-event prediction tasks and the Metaculus benchmark despite a 7x parameter disadvantage.
Benjamin Turtel, Paul Wilczewski, Danny Franklin, Kris Skothiem
arXiv:2601.06336 · cs.LG, cs.AI · submitted Jan 9, 2026 · updated Jan 14, 2026
abstract · pdf · html