about
Evaluating Probabilistic Reasoning in LLMs Through Language-Only Decision Tasks (arxiv.org)
1 point by PaulHoule 332 days ago | hide | past | pdf | 1 comment on HN

In plain words: Models pick an option and get feedback like "you earned a token," so they must learn which pays off best without numbers. Most trailed standard decision rules, but one model chose the best option 89.2% of the time, beating bigger models and classic methods.

Abstract · TextBandit: Evaluating Probabilistic Reasoning in LLMs Through Language-Only Decision Tasks

Large language models (LLMs) have shown to be increasingly capable of performing reasoning tasks, but their ability to make sequential decisions under uncertainty only using natural language remains underexplored. We introduce a novel benchmark in which LLMs interact with multi-armed bandit environments using purely textual feedback, "you earned a token", without access to numerical cues or explicit probabilities, resulting in the model to infer latent reward structures purely off linguistic cues and to adapt accordingly. We evaluated the performance of four open-source LLMs and compare their performance to standard decision-making algorithms such as Thompson Sampling, Epsilon Greedy, Upper Confidence Bound (UCB), and random choice. While most of the LLMs underperformed compared to the baselines, Qwen3-4B, achieved the best-arm selection rate of 89.2% , which significantly outperformed both the larger LLMs and traditional methods. Our findings suggest that probabilistic reasoning is able to emerge from language alone, and we present this benchmark as a step towards evaluating decision-making capabilities in naturalistic, non-numeric contexts.

Jimin Lim, Arjun Damerla, Arthur Jiang, Nam Le
arXiv:2510.13878 · cs.CL · submitted Oct 13, 2025
abstract · pdf · html · COLM 2025 @ ORIGen Workshop

add comment on HN

I see problems:

The paper claims that Qwen3-4B achieved 89.2% best-arm selection by demonstrating superior "probabilistic reasoning". But this is a 2-armed bandit where random guessing should converge to ~50% over 500 runs of 25 iterations each. An 89% rate is suspiciously high and suggests to me that something else is happening (like prompt bias or the model pattern-matching rather than reasoning)

When they increase from 2 to 5 arms, Qwen3-4B drops from 89% to 6.5% accuracy. I assert that if it truly had probabilistic reasoning capability, performance would degrade more gracefully.

The "overthinking" explanation is hand-wavy. I don't see evidence or chain of reasoning. This is just a post-hoc story to explain unexpected results.

No discussion of variance, confidence intervals, or statistical significance. With 500 runs, these should be straightforward to calculate.

Does the claimed 89% accuracy in a binary choice task strike anyone else as implausibly high for what they're claiming?