In plain words: An AI player for the hidden-role game Avalon mixes self-play learning with a planning technique that deduces who is likely friend or foe from partial clues. It beat other programmed and learned players, and outplayed humans both as a teammate and an opponent.
Abstract
Recent breakthroughs in AI for multi-agent games like Go, Poker, and Dota, have seen great strides in recent years. Yet none of these games address the real-life challenge of cooperation in the presence of unknown and uncertain teammates. This challenge is a key game mechanism in hidden role games. Here we develop the DeepRole algorithm, a multi-agent reinforcement learning agent that we test on The Resistance: Avalon, the most popular hidden role game. DeepRole combines counterfactual regret minimization (CFR) with deep value networks trained through self-play. Our algorithm integrates deductive reasoning into vector-form CFR to reason about joint beliefs and deduce partially observable actions. We augment deep value networks with constraints that yield interpretable representations of win probabilities. These innovations enable DeepRole to scale to the full Avalon game. Empirical game-theoretic methods show that DeepRole outperforms other hand-crafted and learned agents in five-player Avalon. DeepRole played with and against human players on the web in hybrid human-agent teams. We find that DeepRole outperforms human players as both a cooperator and a competitor.
Jack Serrino, Max Kleiman-Weiner, David C. Parkes, Joshua B. Tenenbaum
arXiv:1906.02330 · cs.LG, cs.MA, stat.ML · submitted Jun 5, 2019
abstract · pdf · html · Jack Serrino and Max Kleiman-Weiner contributed equally
In the 2189 mixed human/agent games we collected, all humans knew which players were human and which were DeepRole. There were no restrictions on chat usage for the human players, but DeepRole did not say anything and did not process sent messages
I think the fun part about Avalon and other hidden role games really comes from the "cheap talk", where people try to convince each other that their picks / approvals make sense as a member of the good team, as opposed to making the decisions from picks and approvals alone. Though from the results it seems that those concrete actions are already enough to outperform the humans.
There's also the consideration that because the humans know the identity of DeepRole as a bot, they play differently: "That's what I'm gonna pick, because that's what bots do" [1]. I wonder if a combined DeepRole + human-for-chatting-only team would outperform either alone.
[1] From Appendix F of the paper: https://www.youtube.com/watch?v=9RkUFHYTo_s