In plain words: The agent learns how its world works, imagines possible futures, and feeds those predictions to its decision network, which learns how to use them instead of following a planning recipe. It learned faster and stayed reliable even when its guesses about the world were wrong.
Abstract
We introduce Imagination-Augmented Agents (I2As), a novel architecture for deep reinforcement learning combining model-free and model-based aspects. In contrast to most existing model-based reinforcement learning and planning methods, which prescribe how a model should be used to arrive at a policy, I2As learn to interpret predictions from a learned environment model to construct implicit plans in arbitrary ways, by using the predictions as additional context in deep policy networks. I2As show improved data efficiency, performance, and robustness to model misspecification compared to several baselines.
Théophane Weber, Sébastien Racanière, David P. Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adria Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, Razvan Pascanu, Peter Battaglia, et al.
arXiv:1707.06203 · cs.LG, cs.AI, stat.ML · submitted Jul 19, 2017 · updated Feb 14, 2018
abstract · pdf · html