about
Token Reduction Is Not Cost Reduction (arxiv.org)
3 points by umitkaanusta 63 days ago | hide | past | pdf | 1 comment on HN

In plain words: They tested three ways to shrink the text coding assistants read back from tools, then tracked real billed cost, not just token counts. The biggest cut removed 38.4% of tokens yet cost 6.8% more, since caching and extra agent steps ate the savings.

Abstract

Token-reduction tools for coding agents are often evaluated by the number of tokens they remove, but token count alone does not determine end-to-end inference cost. We evaluate three token-reduction approaches against an unmodified Claude Code baseline across controlled coding tasks, measuring provider-billed cost, task success, cache traffic, and agent behavior. The largest compression setup reduced delivered tool-output tokens by 38.4% but increased billed cost by 6.8%, while lighter compression produced only small and statistically uncertain savings. Across tasks, token reduction was weakly correlated with cost reduction (Pearson r = 0.15). Cost decomposition shows that prompt-cache creation and reads dominate the measured input-side cost, leaving only a limited fraction of total spend directly addressable by tool-output compression. We also find that compression can alter agent trajectories through additional retrieval, diagnosis, testing, and turns, offsetting local token savings. On a SWE-bench Go subset, aggressive compression also reduced successful patch application. These results show that token reduction is not a reliable proxy for cost reduction in tool-heavy coding agents. Effective optimization should therefore be evaluated at the level of cost per successful task, including cache behavior, trajectory changes, and correctness rather than token counts alone.

Sarel Weinberger, Amir Hozez
arXiv:2607.12161 · cs.CL · submitted Jul 13, 2026 · updated Aug 12, 2026
abstract · pdf · html

add comment on HN

> Third, aggressive compression can remove action-critical evidence: on SWE-bench-derived Go tasks, compression reduced successful patch application from 27/40 to 15/40 by corrupting verbatim edit anchors.

The only way to avoid this is to hand-craft the translation layer (harness). You cannot rely on some general purpose technique to satisfy domain-specific concerns.

The most extreme version of this I've seen so far is in browser automation. "Compressing" the raw DOM down to a plaintext description (which was parsed deterministically by human-authored code) is a completely different universe of performance than trying to fight the raw DOM directly. Purpose-built compressors can massively outperform generic ones. I've got a case where ~500kb of DOM converts into maybe 1kb of plaintext. There is absolutely no "loss" in terms of information the business cares about in this context. These kinds of decisions cannot be made for you if they are to be made efficiently.