In plain words: The agent learns to predict how the game screen changes after each move, then practices in its own imagined version of the game. With just two hours of real play, it beat trial-and-error methods on most games, sometimes by more than ten times.
Abstract
Model-free reinforcement learning (RL) can be used to learn effective policies for complex tasks, such as Atari games, even from image observations. However, this typically requires very large amounts of interaction -- substantially more, in fact, than a human would need to learn the same games. How can people learn so quickly? Part of the answer may be that people can learn how the game works and predict which actions will lead to desirable outcomes. In this paper, we explore how video prediction models can similarly enable agents to solve Atari games with fewer interactions than model-free methods. We describe Simulated Policy Learning (SimPLe), a complete model-based deep RL algorithm based on video prediction models and present a comparison of several model architectures, including a novel architecture that yields the best results in our setting. Our experiments evaluate SimPLe on a range of Atari games in low data regime of 100k interactions between the agent and the environment, which corresponds to two hours of real-time play. In most games SimPLe outperforms state-of-the-art model-free algorithms, in some games by over an order of magnitude.
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, Afroz Mohiuddin, Ryan Sepassi, et al.
arXiv:1903.00374 · cs.LG, stat.ML · submitted Mar 1, 2019 · updated Apr 3, 2024
abstract · pdf · html
Hasselt argues this is because if you need accurate updates to your Q values, you need to trust the learned model you are using for simulated rollouts on the states on which you are sampling. But if your simulation model is trustworthy on these states, it is because it saw a lot of real transitions from these states from the actual environment. But then you might as well just have stored those transitions in a big enough replay buffer and use ordinary Q-learning with experience replay. And this indeed seems to be the case: when you give Rainbow DQN a nice big replay buffer, it is more sample efficient (both real and imagined samples) than SimPLe. Hasselt leaves some wiggle room for learned models to help with action selection and credit assignment, though.
My counterargument to this (supplied with zero evidence of course!) would be that with the right inductive biases, a learned model can generalize quite accurately and with very few seen transitions, and hence be so sample efficient that it would outperform the replay memory approach. I'd imagine that the kinds of inductive biases that are appropriate for a varied meta-environment like Atari are quite general things like 'visually localized objects typically only interact when they approach or touch each other', and 'the arrow keys likely control one localized object'. There are approaches for how to encode such priors; [2] is a good survey paper, and [3] employs some of these ideas for RL. Moreover, these are the kinds of priors that one imagines are encoded or biased towards by evolution in actual animal brains.
[1] https://arxiv.org/abs/1906.05243
[2] https://arxiv.org/abs/1806.01261
[3] https://arxiv.org/abs/1806.01830