about
Multi-agent cooperation through in-context co-player inference (arxiv.org)
1 point by simonpure 203 days ago | hide | past | pdf | discuss on HN

In plain words: Agents built from sequence models are trained against many different partners, so they learn to adapt to each partner's behavior within one episode instead of assuming a fixed learning rule. This setup naturally produces mutual cooperation, the same shaping trick earlier work engineered by hand.

Abstract

Achieving cooperation among self-interested agents remains a fundamental challenge in multi-agent reinforcement learning. Recent work showed that mutual cooperation can be induced between "learning-aware" agents that account for and shape the learning dynamics of their co-players. However, existing approaches typically rely on hardcoded, often inconsistent, assumptions about co-player learning rules or enforce a strict separation between "naive learners" updating on fast timescales and "meta-learners" observing these updates. Here, we demonstrate that the in-context learning capabilities of sequence models allow for co-player learning awareness without requiring hardcoded assumptions or explicit timescale separation. We show that training sequence model agents against a diverse distribution of co-players naturally induces in-context best-response strategies, effectively functioning as learning algorithms on the fast intra-episode timescale. We find that the cooperative mechanism identified in prior work-where vulnerability to extortion drives mutual shaping-emerges naturally in this setting: in-context adaptation renders agents vulnerable to extortion, and the resulting mutual pressure to shape the opponent's in-context learning dynamics resolves into the learning of cooperative behavior. Our results suggest that standard decentralized reinforcement learning on sequence models combined with co-player diversity provides a scalable path to learning cooperative behaviors.

Marissa A. Weis, Maciej Wołczyk, Rajai Nasser, Rif A. Saurous, Blaise Agüera y Arcas, João Sacramento, Alexander Meulemans
arXiv:2602.16301 · cs.AI · submitted Feb 18, 2026
abstract · pdf · html · 26 pages, 4 figures

add comment on HN
Also discussed: Feb 2026 (2 points, 0 comments)