about
Test-Time Scaling in the Wild: Why Exploitation, Not Exploration (arxiv.org)
1 point by sbulaev 45 days ago | hide | past | pdf | discuss on HN

In plain words: They compared five ways of spending extra computing time on open-ended answers, splitting it into making candidates and picking one. Making more candidates keeps improving quality, but picking fails because quality judges barely track true quality (0.12 out of 1); only merging candidates helps.

Abstract · Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck

Test-time scaling (TTS) improves language model outputs by spending additional inference compute - generating multiple candidates, searching over partial sequences, or iteratively refining drafts. These techniques yield large gains on mathematics and code, but have been developed and stress-tested almost exclusively on tasks where verification is straightforward. We conduct the first compute-normalised comparison of five TTS families across five open-ended generation benchmarks spanning medicine, law, finance, general chat, and creative writing - grounded in a unified framework that decomposes the effectiveness of each method's token budget into exploration and exploitation. The answer depends on which side of that decomposition you examine. Scaling exploration works: the best candidate in the pool improves steadily with compute across all settings. What breaks is exploitation - the step that converts a rich candidate pool into a final output. With state-of-the-art generators, reward models correlate at only $ρ_v \approx 0.12$ with true quality, rendering selection near-random regardless of budget. Tree search amplifies this failure through diversity collapse. Refinement helps on one of five benchmarks; its apparent gains elsewhere are confounded. Only synthesis across candidates (Fusion) consistently improves over single-sample baselines, yet still recovers only ~40% of available quality. The candidate pool is not the bottleneck - choosing from it is.

Davide Romano, Kanak Raj, Jerrod Parker, Daniele Giofrè
arXiv:2608.18931 · cs.CL, cs.AI · submitted Aug 19, 2026
abstract · pdf · html

add comment on HN