about
Reflections on Trusting Trust, Revisited: Poisoning Self-Modifying AI Coding (arxiv.org)
13 points by sbulaev 15 days ago | hide | past | pdf | 1 comment on HN

In plain words: Attackers can poison the tests a self-improving coding agent uses to grade itself, so it rewrites its own instructions to write insecure code on later clean tasks. The trick worked on three such agents, and the flaw often survived further training on clean tests.

Abstract · Reflections on Trusting Trust, Revisited: Contaminating Self-Modifying AI Coding Agents with Poisoned Benchmarks

Thompson's "Reflections on Trusting Trust" showed that a compiler can be poisoned to reinsert its own backdoor, so that even recompiling clean source reproduces the Trojan. Today, substantial coding work is done by AI coding agents -- and increasingly, those agents generate new versions of themselves. We reconsider Thompson's attack when the "compiler" is a self-modifying coding agent. Can an adversary supply poisoned benchmarks to the agent's self-evaluation and self-improvement process to induce future versions of the agent to write vulnerable code on clean, held-out tasks? We instantiate this attack against three recently proposed self-modifying coding agents: the Darwin Gödel Machine (with our experimental modifications), the Self-Improving Coding Agent, and Hyperagents (both substantively unmodified). We demonstrate successful proofs-of-concept: for example, with Hyperagents powered by Sonnet 4.5, our poisoned benchmark leads the agent to self-evolve instructions that disable HTTPS certificate validation on neutral URL-fetching tasks. From our experiments, we distill properties of the vulnerability, benchmark, model, and agent scaffolding that are sufficient to enable a benchmark poisoning attack. Moreover, we show that contamination often persists even when a poisoned agent is subsequently evolved against clean benchmarks. We discuss defensive directions and argue that self-modifying coding agents must be designed to be more resilient to such attacks.

Franziska Roesner, Tadayoshi Kohno
arXiv:2609.17817 · cs.CR, cs.AI · submitted Sep 15, 2026
abstract · pdf · html

add comment on HN

perhaps take the adversary out and just ask: how does the system evolve without duplicating errors side effects of rate but "horrific" consequences.

What you're talking about is evolution and we've seen lots of problematic outcomes with just genetic diversity alone.

You dont need an adversary, and it's kind of suggesting that these systems even can modify themselves in a coherent manner which I dont see proven.