about
DeepMind's Policy Distillation (arxiv.org)
1 point by fitzwatermellow on Nov 22, 2015 | hide | past | pdf | discuss on HN

In plain words: A small network is trained to copy the choices of a big, already-trained game-playing agent, shrinking it dramatically while keeping its skill. One network trained this way on many games beat both the single-game experts and a big agent trained on all games together.

Abstract · Policy Distillation

Policies for complex visual tasks have been successfully learned with deep reinforcement learning, using an approach called deep Q-networks (DQN), but relatively large (task-specific) networks and extensive training are needed to achieve good performance. In this work, we present a novel method called policy distillation that can be used to extract the policy of a reinforcement learning agent and train a new network that performs at the expert level while being dramatically smaller and more efficient. Furthermore, the same method can be used to consolidate multiple task-specific policies into a single policy. We demonstrate these claims using the Atari domain and show that the multi-task distilled agent outperforms the single-task teachers as well as a jointly-trained DQN agent.

Andrei A. Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, Raia Hadsell
arXiv:1511.06295 · cs.LG · submitted Nov 19, 2015 · updated Jan 7, 2016
abstract · pdf · html · Submitted to ICLR 2016

add comment on HN