In plain words: AlphaZero normally learns two-player games by playing against itself; here it is adjusted to handle three or more players at once. In two simple three-player games, it learned winning strategies and consistently beat the usual tree-search player that guesses ahead by sampling random games.
Abstract
The AlphaZero algorithm has achieved superhuman performance in two-player, deterministic, zero-sum games where perfect information of the game state is available. This success has been demonstrated in Chess, Shogi, and Go where learning occurs solely through self-play. Many real-world applications (e.g., equity trading) require the consideration of a multiplayer environment. In this work, we suggest novel modifications of the AlphaZero algorithm to support multiplayer environments, and evaluate the approach in two simple 3-player games. Our experiments show that multiplayer AlphaZero learns successfully and consistently outperforms a competing approach: Monte Carlo tree search. These results suggest that our modified AlphaZero can learn effective strategies in multiplayer game scenarios. Our work supports the use of AlphaZero in multiplayer games and suggests future research for more complex environments.
Nick Petosa, Tucker Balch
arXiv:1910.13012 · cs.AI · submitted Oct 29, 2019 · updated Dec 9, 2019
abstract · pdf · html
The comparison against MCTS shows strong performance from AlphaZero. Would be curious to see the performance of AlphaZero vs. number of its own rollouts - ie is the probability head output alone already encoding enough information to play well, and how deep does it have to look ahead to determine good play. Finally, for tic-tac-mo and connect 3x3, it should be possible to determine the optimal move. How much training / lookahead is required to achieve that? Does AlphaZero achieve perfect play for these games?
The paper's first listed contribution is "an independent reimplementation of DeepMind's AlphaZero algorithm". Maybe I missed it, but I don't see a link to a repo with the implementation.