about
Latent Programming Horizons in Coding Agents (arxiv.org)
3 points by andre15silva 88 days ago | hide | past | pdf | 2 comments on HN

In plain words: A simple classifier reading a coding agent's internal states can tell whether its code parses and passes tests, and guess what future edits will do. It reached 0.83 for correctness and stayed useful up to 25 steps ahead, on new tasks without retraining.

Abstract

A coding agent solving a software-engineering task spends dozens of steps reasoning, editing code, and running tests, yet little is known about what the underlying language model internally represents about the program it is working on. We show that the residual streams of language models under coding agents linearly encode properties of the evolving program: a logistic-regression probe on hidden states is able to decode whether the current code parses, passes its test suite, reduces the number of failing tests, and introduces regressions, reaching AUC up to 0.83 for correctness across two models and two benchmarks. Our second finding is more surprising: these representations run ahead of the agent's own edits. Probes trained to predict the outcome of future edits (before they are materialized and written on disk) achieve performance above chance up to roughly 25 steps in advance. We call this the agent's latent programming horizon. As a proof of external validity, we show that the probes transfer across benchmarks without retraining. Our positive results open calls for more research in mechanistic interpretability of coding agents.

André Silva, Han Tu, Martin Monperrus
arXiv:2607.05188 · cs.LG, cs.SE · submitted Jul 6, 2026
abstract · pdf · html

add comment on HN
Also discussed: Jul 2026 (96 points, 78 comments)

Nice paper. They predict the outcome of edits up to ~25 steps before the agent makes them. Decodable doesn't mean causal, as the authors note, but a cheap probe that flags doomed trajectories early could save a lot of wasted agent compute. We clearly still have a lot to learn about what these models represent internally.
good to read