about
Neural Episodic Control (arxiv.org)
50 points by beefman on Mar 11, 2017 | hide | past | pdf | 5 comments on HN

In plain words: The agent keeps a memory of past situations, with slowly changing descriptions but instantly updated guesses of how good each one is, so a new experience changes its next move right away. Across many game environments it learned far faster than other general-purpose agents.

Abstract

Deep reinforcement learning methods attain super-human performance in a wide range of environments. Such methods are grossly inefficient, often taking orders of magnitudes more data than humans to achieve reasonable performance. We propose Neural Episodic Control: a deep reinforcement learning agent that is able to rapidly assimilate new experiences and act upon them. Our agent uses a semi-tabular representation of the value function: a buffer of past experience containing slowly changing state representations and rapidly updated estimates of the value function. We show across a wide range of environments that our agent learns significantly faster than other state-of-the-art, general purpose deep reinforcement learning agents.

Alexander Pritzel, Benigno Uria, Sriram Srinivasan, Adrià Puigdomènech, Oriol Vinyals, Demis Hassabis, Daan Wierstra, Charles Blundell
arXiv:1703.01988 · cs.LG, stat.ML · submitted Mar 6, 2017
abstract · pdf · html

add comment on HN
Also discussed: Mar 2017 (4 points, 0 comments)

I wonder how they deal with the change of the mapping from states to vectors h from the CNN changing as training advances.

I reckon they _don't_ train the CNN, rather, they use it as pre-trained from DQN or something. In that case, no wonder it learns faster.

they do train the CNN from scratch, see section 3.4.
Note that Denis Hassibis is one of the authors and that all authors are at Google, Deep Mind.

Without this I'd be quite skeptical of these claims but now I'm wishing for analysis from someone better qualified than I am.

A glance of the paper seems to suggest that they are going back to the traditional RL methods. It would be interesting for the paper to compare with them explicitly.
Which traditional RL methods do you mean, and what kind of comparison?

If you mean they are returning to tabular Q-learning, then yes, sort of. They still use function approximation (as required; tabular Q-learning stands no chance of generalising across states), but their function kind of looks like a table look up. A look up of the 50 closest values to some computed key though.

It would be totally uninteresting to compare this with tabular Q-learning though.