about
Unsupervised Reinforcement Learning for Intrinsic Motivation (arxiv.org)
3 points by AndrewKemendo on Nov 30, 2016 | hide | past | pdf | discuss on HN

In plain words: An agent with no teacher learns skills by trying to reach as many different ending states as possible, so each skill reliably leads somewhere distinct. It scales to large networks, works on several tasks, and scores how much control a state offers.

Abstract · Variational Intrinsic Control

In this paper we introduce a new unsupervised reinforcement learning method for discovering the set of intrinsic options available to an agent. This set is learned by maximizing the number of different states an agent can reliably reach, as measured by the mutual information between the set of options and option termination states. To this end, we instantiate two policy gradient based algorithms, one that creates an explicit embedding space of options and one that represents options implicitly. The algorithms also provide an explicit measure of empowerment in a given state that can be used by an empowerment maximizing agent. The algorithm scales well with function approximation and we demonstrate the applicability of the algorithm on a range of tasks.

Karol Gregor, Danilo Jimenez Rezende, Daan Wierstra
arXiv:1611.07507 · cs.LG, cs.AI · submitted Nov 22, 2016
abstract · pdf · html · 15 pages, 6 figures

add comment on HN