In plain words: Chatbots played repeated two-player games against each other and real people to test how well they cooperate and coordinate. They excelled at self-interested games like the Prisoner's Dilemma but struggled at coordination games; telling GPT-4 about its opponent raised scores and teamwork with humans.
Abstract · Playing repeated games with Large Language Models
LLMs are increasingly used in applications where they interact with humans and other agents. We propose to use behavioural game theory to study LLM's cooperation and coordination behaviour. We let different LLMs play finitely repeated $2\times2$ games with each other, with human-like strategies, and actual human players. Our results show that LLMs perform particularly well at self-interested games like the iterated Prisoner's Dilemma family. However, they behave sub-optimally in games that require coordination, like the Battle of the Sexes. We verify that these behavioural signatures are stable across robustness checks. We additionally show how GPT-4's behaviour can be modulated by providing additional information about its opponent and by using a "social chain-of-thought" (SCoT) strategy. This also leads to better scores and more successful coordination when interacting with human players. These results enrich our understanding of LLM's social behaviour and pave the way for a behavioural game theory for machines.
Elif Akata, Lion Schulz, Julian Coda-Forno, Seong Joon Oh, Matthias Bethge, Eric Schulz
arXiv:2305.16867 · cs.CL · submitted May 26, 2023 · updated May 7, 2025
abstract · pdf · html