In plain words: A neural network learns to play Counter-Strike from raw screen pixels by copying millions of frames of real human matches, mixing noisy online games with a few expert demos. It matches the game's medium-difficulty built-in bot in deathmatch while playing in a human style.
Abstract · Counter-Strike Deathmatch with Large-Scale Behavioural Cloning
This paper describes an AI agent that plays the popular first-person-shooter (FPS) video game `Counter-Strike; Global Offensive' (CSGO) from pixel input. The agent, a deep neural network, matches the performance of the medium difficulty built-in AI on the deathmatch game mode, whilst adopting a humanlike play style. Unlike much prior work in games, no API is available for CSGO, so algorithms must train and run in real-time. This limits the quantity of on-policy data that can be generated, precluding many reinforcement learning algorithms. Our solution uses behavioural cloning - training on a large noisy dataset scraped from human play on online servers (4 million frames, comparable in size to ImageNet), and a smaller dataset of high-quality expert demonstrations. This scale is an order of magnitude larger than prior work on imitation learning in FPS games.
Tim Pearce, Jun Zhu
arXiv:2104.04258 · cs.AI, cs.LG, stat.ML · submitted Apr 9, 2021 · updated Dec 9, 2021
abstract · pdf · html · Offline Reinforcement Learning Workshop at Neural Information Processing Systems, 2021
Finally, you just do dead-simple behavior cloning (predict "expert" action from current observation) on the large labeled dataset. It's still surprising to me how well this works! Behavior cloning has some theoretical issues since it doesn't address the sequential nature of decision problems at all, but apparently with enough scale it can still do fairly well - super cool.
[1]: https://openai.com/blog/vpt/