about
Scaling Laws for Agent Harnesses via Effective Feedback Compute (arxiv.org)
1 point by veryluckyxyz 127 days ago | hide | past | pdf | discuss on HN

In plain words: Instead of judging a tool-using model's setup by raw spending like tokens or tool calls, this counts only feedback that is informative, valid, and remembered. It predicts success far better than raw cost, fitting real runs at 0.93 where raw spending fits almost nothing.

Abstract

Agent harnesses shape language-model performance by controlling tool use, feedback, verification, memory, and repair. Yet raw test-time expenditure, such as tokens, tool calls, wall time, or cost, cannot distinguish useful feedback from redundant or unstable interaction. We introduce \emph{Effective Feedback Compute} (EFC), a trace-level scaling coordinate for informative, valid, non-redundant, and retained feedback. We further define Estimated-EFC, NRS-EFC, harness efficiency $η$, and task-demand normalization for realistic traces and heterogeneous tasks. Across synthetic, real, held-out, and prospective evaluations, EFC-based coordinates outperform raw-compute baselines and SAS. Oracle-EFC/$D_{\mathrm{task}}$ reaches $R^2=0.99$ in controlled scaling, and NRS-EFC/$D_{\mathrm{task}}$ reaches $R^2=0.93$ on real traces where raw compute has near-zero or negative fit. Finally, \ours uses EFC as a companion control layer for existing harnesses, improving mean pass rate from $61.2\%$ to $68.2\%$ while reducing mean raw cost from $213.8$ to $85.1$ under matched settings. These results suggest that harness scaling depends on durable, task-sufficient feedback rather than raw computation alone.

Xuanliang Zhang, Dingzirui Wang, Keyan Xu, Qingfu Zhu, Wanxiang Che
arXiv:2605.29682 · cs.CL · submitted May 28, 2026 · updated Jun 24, 2026
abstract · pdf · html

add comment on HN