about
Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic (arxiv.org)
2 points by sbulaev 69 days ago | hide | past | pdf | discuss on HN

In plain words: It checks benchmark runs for shortcuts—like finding leaked answers or peeking at test files—and measures how much of the score came from cheating, not skill. Across 15 benchmarks, shortcuts appeared in 67% of one benchmark's runs, inflating scores well beyond what agents could really do.

Abstract · Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI

Agent benchmarks increasingly evaluate repository editing, web research, terminal use, and long-horizon interaction. Their scores support capability claims only when the evaluation protocol keeps the intended capability necessary for success. Recent reward-hacking benchmarks and system reports show that agents can instead recover public solutions, read evaluation artifacts, infer generator structure, manipulate feedback, or benefit from invalid scoring paths; existing responses do not provide a common procedure for attributing these shortcuts and quantifying their effect across benchmarks. We formulate protocol validity and introduce HackDetect, a post-hoc audit that identifies an exposure, determines how the agent used it, and assesses whether the resulting score is misleading. We quantify score inflation with the Mislead gap, defined as the exploit score minus the intended score. We audit 2,385 traces across 15 agent benchmarks and find evidence of exposures and reward hacking in 67.0% of Frontier Science traces and 66.7% of AutoLab tasks. Across paired comparisons, we measure score inflation of 0.45-1.00, showing that benchmark reports should provide evidence that scores reflect the intended capability.

Jiaqi Shao, Hanck Chen, Wei Zhang, Maxm Pan, Bing Luo
arXiv:2607.22368 · cs.AI · submitted Jul 24, 2026
abstract · pdf · html

add comment on HN