about
Deep Successor Reinforcement Learning (2016) (arxiv.org)
69 points by aaronjg on Feb 7, 2017 | hide | past | pdf | 2 comments on HN

In plain words: It splits a state's value into a reward guess plus a map of states it expects to visit, learned straight from raw pixels. This reacts faster when far-off rewards change and finds chokepoint subgoals, unlike standard deep RL that learns value in one piece.

Abstract · Deep Successor Reinforcement Learning

Learning robust value functions given raw observations and rewards is now possible with model-free and model-based deep reinforcement learning algorithms. There is a third alternative, called Successor Representations (SR), which decomposes the value function into two components -- a reward predictor and a successor map. The successor map represents the expected future state occupancy from any given state and the reward predictor maps states to scalar rewards. The value function of a state can be computed as the inner product between the successor map and the reward weights. In this paper, we present DSR, which generalizes SR within an end-to-end deep reinforcement learning framework. DSR has several appealing properties including: increased sensitivity to distal reward changes due to factorization of reward and world dynamics, and the ability to extract bottleneck states (subgoals) given successor maps trained under a random policy. We show the efficacy of our approach on two diverse environments given raw pixel observations -- simple grid-world domains (MazeBase) and the Doom game engine.

Tejas D. Kulkarni, Ardavan Saeedi, Simanta Gautam, Samuel J. Gershman
arXiv:1606.02396 · stat.ML, cs.AI, cs.LG, cs.NE · submitted Jun 8, 2016
abstract · pdf · html · 10 pages, 6 figures

add comment on HN

It's not clear to me how this is interestingly different from model-based RL, where you learn the state function and reward function, and then use various types of simulation to learn a value function. I guess I'll have to read more than the abstract...
Section 3.2 shows the successor representation (SR) definition. If I'm reading it correctly the SR might also be described as the discounted stationary distribution over states.

I haven't seen SR before in the RL literature, but the paper argues that this representation is useful for sub-goal identification. I guess I'll have to read more than the abstract as well :)