about
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering (arxiv.org)
2 points by wek 98 days ago | hide | past | pdf | discuss on HN

In plain words: Today's coding benchmarks grade a whole agent setup as one score against a single reference answer, so they can't show which piece helped or accept other valid solutions. Swapping any one piece can move the score as much as upgrading to a newer model.

Abstract

Coding agents have become a major mode of software engineering, but the benchmarks we use to compare them were designed in a pre-agent era: they collapse model, harness, and environment into a single end-to-end score, typically computed against one reference solution, with no component-level signal for iteration. We argue that current coding benchmarks are misaligned with agentic software engineering. A coding agent in practice is not a model: it is a system harness -- a composite of models, harnesses, contexts, environments, and feedback signals, any one of which can move the benchmark score by margins comparable to those between adjacent model generations. We discuss three symptoms: (i) benchmark scores conflate the model with the rest of the harness; (ii) grading against a single reference solution penalises equally valid alternatives; and (iii) the absence of signal at the level of individual harness components makes the end-to-end system score difficult to iterate on.

Maria I. Gorinova, Macey Baker, Amy Heineike, Maksim Shaposhnikov, Rob Willoughby, Dru Knox
arXiv:2606.17799 · cs.SE, cs.AI, cs.CL · submitted Jun 16, 2026 · updated Jul 18, 2026
abstract · pdf · html

add comment on HN
Also discussed: Jun 2026 (1 point, 1 comment)