In plain words: The model writes its own guesses and reasoning for questions about the future, then gets trained to prefer the guesses that later matched what really happened. This raised accuracy by 7–10% over the untrained and randomly trained versions, matching much larger models.
Abstract · LLMs Can Teach Themselves to Better Predict the Future
We present an outcome-driven fine-tuning framework that enhances the forecasting capabilities of large language models (LLMs) without relying on human-curated reasoning samples. Our method leverages model self-play to generate pairs of diverse reasoning trajectories and probabilistic forecasts for a set of diverse questions that resolve after the models' knowledge cutoff date. We then rank pairs of these reasoning traces by their distance to the actual outcomes before fine-tuning the model via Direct Preference Optimization (DPO). On a separate test set, our approach increases prediction accuracy of Phi-4 14B and DeepSeek-R1 14B by between 7--10\% over a base model and a DPO fine-tuned control model with randomized labels, bringing them on par with forecasting capabilities of much larger frontier models like GPT-4o.
Benjamin Turtel, Danny Franklin, Philipp Schoenegger
arXiv:2502.05253 · cs.CL, cs.AI · submitted Feb 7, 2025
abstract · pdf · html
... [T]hese researchers are working long hours to put themselves out of a job. They need AI agents that can think ahead, so engineers train agents to forecast. They hold out training data before 2024, instructing models to ponder for hours to predict events in 2025. Then, they apply the same trick as before, distilling pondering into a gut reaction. Forecasting ability is a broad foundation. The researchers build specialized ML research skills on top of it, training U3 to predict the results of every ML paper and ML experiment ever recorded.
[0] https://www.lesswrong.com/posts/KFJ2LFogYqzfGB3uX/how-ai-tak...
[1] https://news.ycombinator.com/item?id=43004579