about
Forecasting Ability of LLMs Depends on What We're Asking (arxiv.org)
1 point by paraschopra 312 days ago | hide | past | pdf | 1 comment on HN

In plain words: They tested several AI chat models on real-world questions about events that happened after the models' training ended, changing the topic, question wording, and news given. Accuracy swung widely depending on what was asked and how, so no single score describes their forecasting skill.

Abstract · Future Is Unevenly Distributed: Forecasting Ability of LLMs Depends on What We're Asking

Large Language Models (LLMs) demonstrate partial forecasting competence across social, political, and economic events. Yet, their predictive ability varies sharply with domain structure and prompt framing. We investigate how forecasting performance varies with different model families on real-world questions about events that happened beyond the model cutoff date. We analyze how context, question type, and external knowledge affect accuracy and calibration, and how adding factual news context modifies belief formation and failure modes. Our results show that forecasting ability is highly variable as it depends on what, and how, we ask.

Chinmay Karkar, Paras Chopra
arXiv:2511.18394 · cs.LG · submitted Nov 23, 2025
abstract · pdf · html

add comment on HN

>> Our results show that forecasting ability is highly variable as it depends on what, and how, we ask.

Do we really call this "research"?