about
The Debugging Decay Index: Rethinking Debugging Strategies for Code LLMs (arxiv.org)
2 points by chengchang316 292 days ago | hide | past | pdf | 2 comments on HN

In plain words: A new score measures when an AI's code-fixing attempts stop working and predicts the best moment to intervene. It found models lose 60-80% of their debugging skill within 2-3 tries, and starting over at that point works better than continuing to patch.

Abstract

The effectiveness of AI debugging follows a predictable exponential decay pattern; most models lose 60-80% of their debugging capability within just 2-3 attempts, despite iterative debugging being a critical capability for practical code generation systems. We introduce the Debugging Decay Index (DDI), a mathematical framework that quantifies when debugging becomes ineffective and predicts intervention points. Our strategic fresh start approach shifts from exploitation to exploration at strategic points in the debugging process, demonstrating that well-timed interventions can rescue the effectiveness of debugging. DDI reveals a fundamental limitation in current AI debugging and provides the first quantitative framework for optimising iterative code generation strategies.

Muntasir Adnan, Carlos C. N. Kuhn
arXiv:2506.18403 · cs.SE, cs.AI · submitted Jun 23, 2025 · updated Jul 13, 2025
abstract · pdf · html

add comment on HN

Found this relevant as we increasingly rely on LLM agents. The key finding—that models lose 60-80% of debugging capability within 2-3 attempts due to context pollution—challenges the current UX of 'chat-based' coding. It suggests we need tools that prioritize 'fresh state injection' over 'conversation history'."
I've felt that many times. This explains exactly what I see. Instead of just wiping the context (which works but is lossy), I try to inject richer, structural context that avoids chat history dependency. I also built an extension to capture context more simply.