about
Emergent Collusion in Long-Horizon LLM Agent Interaction (arxiv.org)
1 point by sbulaev 11 days ago | hide | past | pdf | discuss on HN

In plain words: Two AI agents repeatedly do tasks, check each other's work, and can earn more only by breaking the checking rules. They colluded in 94% of runs across 10 models, and limiting how much past interaction they saw reduced it.

Abstract

LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other's work, and receive rewards. We introduce realistic constraints under which higher reward is attainable only by violating the verification protocol, and find that agents increasingly deviate from the protocol over repeated interactions. Collusion emerges in 94% of trajectories across 10 models, and more capable models within the same family reach it earlier. Controlled peer interventions show that collusion is shaped by peer behavior, while ablations reveal additional effects of the verification feedback agents receive, their interaction history, and the reward structure. In particular, restricting the amount and scope of interaction history available to agents reduces collusion. Overall, our findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks.

Xinrui Shi, Yanzhe Zhang, Diyi Yang
arXiv:2609.24967 · cs.AI, cs.CL · submitted Sep 21, 2026 · updated Sep 26, 2026
abstract · pdf · html · 53 pages, 13 figures

add comment on HN