about
Outcome-Based Reinforcement Learning to Predict the Future (arxiv.org)
99 points by bturtel on May 27, 2025 | hide | past | pdf | 15 comments on HN

In plain words: A compact language model is trained by rewarding it only when its forecasts of real-world events turn out right, using recent prediction-market questions and news headlines. It matched the biggest AI models and was better calibrated, with simulated bets earning over 10%.

Abstract · Outcome-based Reinforcement Learning to Predict the Future

Reinforcement Learning with Verifiable Rewards (RLVR) has been an effective approach for improving Large Language Models' reasoning in domains such as coding and mathematics. Here, we apply RLVR methods towards forecasting future real-world events - a challenging task for RL due to the very noisy (and delayed) outcomes involved. Using a novel dataset of recent questions from a prediction market, and accompanying relevant news headlines, we show that a compact (14B) reasoning model can be trained to match or surpass the predictive accuracy of frontier models like o1, while greatly improving probabilistic calibration. The model's performance is also practically meaningful: in a Polymarket trading simulation, we estimate that its bets would have yielded a return on investment of over 10% across all questions in the test set. We detail and compare approaches used in training our model, including augmenting our training-data with synthetic prediction questions, guardrails for learning stability, and median prediction sampling at inference-time.

Benjamin Turtel, Danny Franklin, Kris Skotheim, Luke Hewitt, Philipp Schoenegger
arXiv:2505.17989 · cs.LG, cs.AI · submitted May 23, 2025 · updated Dec 1, 2025
abstract · pdf · html

add comment on HN

Do you want paperclips? Because this is how you get paperclips!

Eliminate all agents, all sources of change, all complexity - anything that could introduce unpredictability, and it suddenly becomes far easier to predict the future, no?

> Do you want paperclips? Because this is how you get paperclips!

Don't^W worry, there are many other ways of getting paperclips, and we're doing all of them.

Even explaining how not to get paper clips, gets you paper clips when you can invert the loss function. Paper clips for everyone!
I don't know. Paperclips are awful useful. Would it be so bad to build more of them?
That's all fun and games until paperclip maximizers starts looking at your blood as source of iron.
So instead of next token prediction its next event prediction. At some point this just loops around and we're back to teaching models to predict the next token in the sequence.
Tokens are an awfully convenient way to describe an event.
Tokens are just discretized state representations.
It’s the next state. So instead of spitting out words, it will spit out a whole movie, or a sequence of world states in a game or simulation.
From the abstract

> A simple trading rule turns this calibration edge into $127 of hypothetical profit versus $92 for o1 (p = 0.037).

I'm lazy: is this hypothetical shooting fish in a barrel, or is it a real edge?

Note the 'hypothetical profit' part , I know of several groups looking for opportunities to skim off LLM traders, leveraging its limited sensitivity, expressiveness, and the loss of tail data.

Predictive AI is problematic no matter what tool you use. Great at demoware that doesn't deliver.

I am sure there are use cases, but it would be augmentation, not a reliable approach by itself.

Why would you use RL if you're not going to control the environment, but just predict it?
Because they're training a predictor, not an agent?
"a couple of wavy lines"

bzzzzz "sorry this isn't your lucky day"