about
Open-Ended Learning in Symmetric Zero-Sum Games (arxiv.org)
3 points by lawrenceyan on Jan 28, 2019 | hide | past | pdf | discuss on HN

In plain words: In games where strategies loop like rock-paper-scissors, there's no single stronger opponent to chase, so the training goal must keep changing. The rule built here keeps a diverse crowd of strong agents and beat the usual approaches on two resource-allocation games.

Abstract · Open-ended Learning in Symmetric Zero-sum Games

Zero-sum games such as chess and poker are, abstractly, functions that evaluate pairs of agents, for example labeling them `winner' and `loser'. If the game is approximately transitive, then self-play generates sequences of agents of increasing strength. However, nontransitive games, such as rock-paper-scissors, can exhibit strategic cycles, and there is no longer a clear objective -- we want agents to increase in strength, but against whom is unclear. In this paper, we introduce a geometric framework for formulating agent objectives in zero-sum games, in order to construct adaptive sequences of objectives that yield open-ended learning. The framework allows us to reason about population performance in nontransitive games, and enables the development of a new algorithm (rectified Nash response, PSRO_rN) that uses game-theoretic niching to construct diverse populations of effective agents, producing a stronger set of agents than existing algorithms. We apply PSRO_rN to two highly nontransitive resource allocation games and find that PSRO_rN consistently outperforms the existing alternatives.

David Balduzzi, Marta Garnelo, Yoram Bachrach, Wojciech M. Czarnecki, Julien Perolat, Max Jaderberg, Thore Graepel
arXiv:1901.08106 · cs.LG, cs.GT, cs.MA, stat.ML · submitted Jan 23, 2019 · updated May 13, 2019
abstract · pdf · html · ICML 2019, final version

add comment on HN