about
Rethinking the Value of Generated Tests for LLM Software Engineering Agents (arxiv.org)
2 points by zuzululu 120 days ago | hide | past | pdf | discuss on HN

In plain words: They studied how often six coding agents write their own tests while fixing real GitHub bugs, and whether those tests help. Solved and unsolved tasks had similar test-writing rates, and prompting agents to write more or fewer tests didn't change final results.

Abstract · Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agents

Large Language Model (LLM) code agents increasingly resolve repository-level issues by iteratively editing code, invoking tools, and validating candidate patches. In these workflows, agents often write tests on the fly, but the value of this behavior remains unclear. For example, GPT-5.2 writes almost no new tests yet achieves performance comparable to top-ranking agents.This raises a central question: do such tests meaningfully improve issue resolution, or do they mainly mimic a familiar software-development practice while consuming interaction budget? To better understand the role of agent-written tests, we analyze trajectories produced by six strong LLMs on SWE-bench Verified. Our results show that test writing is common, but resolved and unresolved tasks within the same model exhibit similar test-writing frequencies. When tests are written, they mainly serve as observational feedback channels, with value-revealing print statements appearing much more often than assertion-based checks. Based on these insights, we perform a prompt-intervention study by revising the prompts used with four models to either increase or reduce test writing. The results suggest that prompt-induced changes in the volume of agent-written tests do not significantly change final outcomes in this setting. Taken together, these results suggest that current agent-written testing practices reshape process and cost more than final task outcomes.

Zhi Chen, Zhensu Sun, Yuling Shi, Chao Peng, Xiaodong Gu, David Lo, Lingxiao Jiang
arXiv:2602.07900 · cs.SE, cs.AI · submitted Feb 8, 2026 · updated Apr 9, 2026
abstract · pdf · html

add comment on HN