In plain words: The agent keeps a memory of past situations, with slowly changing descriptions but instantly updated guesses of how good each one is, so a new experience changes its next move right away. Across many game environments it learned far faster than other general-purpose agents.
Abstract
Deep reinforcement learning methods attain super-human performance in a wide range of environments. Such methods are grossly inefficient, often taking orders of magnitudes more data than humans to achieve reasonable performance. We propose Neural Episodic Control: a deep reinforcement learning agent that is able to rapidly assimilate new experiences and act upon them. Our agent uses a semi-tabular representation of the value function: a buffer of past experience containing slowly changing state representations and rapidly updated estimates of the value function. We show across a wide range of environments that our agent learns significantly faster than other state-of-the-art, general purpose deep reinforcement learning agents.
Alexander Pritzel, Benigno Uria, Sriram Srinivasan, Adrià Puigdomènech, Oriol Vinyals, Demis Hassabis, Daan Wierstra, Charles Blundell
arXiv:1703.01988 · cs.LG, stat.ML · submitted Mar 6, 2017
abstract · pdf · html
I reckon they _don't_ train the CNN, rather, they use it as pre-trained from DQN or something. In that case, no wonder it learns faster.