about
Can AI Agents Agree? (arxiv.org)
1 point by tanelpoder 209 days ago | hide | past | pdf | discuss on HN

In plain words: Groups of AI agents were simulated trying to agree on a number, with some agents trying to derail the deal. Agreement often failed even with no disruptors and got worse as groups grew, usually by stalling rather than settling on a wrong value.

Abstract

Large language models are increasingly deployed as cooperating agents, yet their behavior in adversarial consensus settings has not been systematically studied. We evaluate LLM-based agents on a Byzantine consensus game over scalar values using a synchronous all-to-all simulation. We test consensus in a no-stake setting where agents have no preferences over the final value, so evaluation focuses on agreement rather than value optimality. Across hundreds of simulations spanning model sizes, group sizes, and Byzantine fractions, we find that valid agreement is not reliable even in benign settings and degrades as group size grows. Introducing a small number of Byzantine agents further reduces success. Failures are dominated by loss of liveness, such as timeouts and stalled convergence, rather than subtle value corruption. Overall, the results suggest that reliable agreement is not yet a dependable emergent capability of current LLM-agent groups even in no-stake settings, raising caution for deployments that rely on robust coordination.

Frédéric Berdoz, Leonardo Rugli, Roger Wattenhofer
arXiv:2603.01213 · cs.MA, cs.LG · submitted Mar 1, 2026 · updated Mar 12, 2026
abstract · pdf · html

add comment on HN