about
RL²: Fast Reinforcement Learning via Slow Reinforcement Learning (2016) (arxiv.org)
50 points by saycheese on Jan 19, 2017 | hide | past | pdf | 9 comments on HN

In plain words: Instead of hand-designing a quick learner, they train a memory network slowly across many tasks, so its stored state acts as a learner that improves on each new task. On fresh tasks it nearly matched hand-built algorithms with guaranteed optimality and handled vision navigation.

Abstract · RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Deep reinforcement learning (deep RL) has been successful in learning sophisticated behaviors automatically; however, the learning process requires a huge number of trials. In contrast, animals can learn new tasks in just a few trials, benefiting from their prior knowledge about the world. This paper seeks to bridge this gap. Rather than designing a "fast" reinforcement learning algorithm, we propose to represent it as a recurrent neural network (RNN) and learn it from data. In our proposed method, RL$^2$, the algorithm is encoded in the weights of the RNN, which are learned slowly through a general-purpose ("slow") RL algorithm. The RNN receives all information a typical RL algorithm would receive, including observations, actions, rewards, and termination flags; and it retains its state across episodes in a given Markov Decision Process (MDP). The activations of the RNN store the state of the "fast" RL algorithm on the current (previously unseen) MDP. We evaluate RL$^2$ experimentally on both small-scale and large-scale problems. On the small-scale side, we train it to solve randomly generated multi-arm bandit problems and finite MDPs. After RL$^2$ is trained, its performance on new MDPs is close to human-designed algorithms with optimality guarantees. On the large-scale side, we test RL$^2$ on a vision-based navigation task and show that it scales up to high-dimensional problems.

Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, Pieter Abbeel
arXiv:1611.02779 · cs.AI, cs.LG, cs.NE, stat.ML · submitted Nov 9, 2016 · updated Nov 10, 2016
abstract · pdf · html · 14 pages. Under review as a conference paper at ICLR 2017

add comment on HN

This seems to be the important bit, that describes what makes their learning "fast":

"The objective (...) is to maximize the expected total discounted reward accumulated during a single trial rather than a single episode. Maximizing this objective is equivalent to minimizing the cumulative pseudo-regret (Bubeck & Cesa-Bianchi, 2012). Since the underlying MDP changes across trials, as long as different strategies are required for different MDPs, the agent must act differently according to its belief over which MDP it is currently in. Hence, the agent is forced to integrate all the information it has received, including past actions, rewards, and termination flags, and adapt its strategy continually. Hence, we have set up an end-to-end optimization process, where the agent is encouraged to learn a “fast” reinforcement learning algorithm"

However, one would note that this is still bounded by learning of the RNN, so I don't really see how this approach makes the algorithm much faster than "slow" RL, as loads of trials would still be required for any real learning. Maybe someone more knowledgeable could pitch in.

(Author here)

This is a great question! There are at least two scenarios where having a meta-learning setup could help. The first is to learn the fast RL algorithm in simulation, by having a distribution over real world environments (varying physics, textures, etc), and then run it in the real world. The second is to learn a fast RL algorithm over a wide range of tasks (for example, on a set of training games in Universe), and (hopefully) generalize to unseen games, analogous to generalization in supervised learning.

Does OpenAI have a page listing all of their published papers?
If this works, this seems like it could be very significant.

Broadly, slowness is a serious problem in current machine learning approaches and so anything that speeds things up is significant.

The approach of learning the learning process would seem to get bonus points for being interesting and general.

Edit: Published in November, this paper didn't seem to get any comments on the machine learning Reddit, which is my go-to for informed on this stuff. I'd love to have someone who knew what they were doing comment here.

Note: The original title noted OpenAI's role in the paper, which maybe seen by loading the PDF and reading the author credits.

https://arxiv.org/pdf/1611.02779v2.pdf

So when are we going to see the paper where we use an RL net to speed up another RL net for discovering a neural architecture for learning how to do gradient descent (by gradient descent)?
Learning to learn by gradient descent by gradient descent: https://arxiv.org/abs/1606.04474
So...you mean a grad student? ;)
It's called Grad student descent (referring to the grad student's descent into despair :P)