about
Yoshua Bengio: Unsupervised Learning of Influential Trajectories (arxiv.org)
3 points by Anon84 on May 24, 2019 | hide | past | pdf | discuss on HN

In plain words: An agent explores with no outside rewards by learning action sequences that maximize how much it changes its environment's future states, judging the whole path instead of just the final state. It works even when the agent has a huge set of possible actions.

Abstract · The Journey is the Reward: Unsupervised Learning of Influential Trajectories

Unsupervised exploration and representation learning become increasingly important when learning in diverse and sparse environments. The information-theoretic principle of empowerment formalizes an unsupervised exploration objective through an agent trying to maximize its influence on the future states of its environment. Previous approaches carry certain limitations in that they either do not employ closed-loop feedback or do not have an internal state. As a consequence, a privileged final state is taken as an influence measure, rather than the full trajectory. We provide a model-free method which takes into account the whole trajectory while still offering the benefits of option-based approaches. We successfully apply our approach to settings with large action spaces, where discovery of meaningful action sequences is particularly difficult.

Jonathan Binas, Sherjil Ozair, Yoshua Bengio
arXiv:1905.09334 · cs.LG, cs.AI, stat.ML · submitted May 22, 2019
abstract · pdf · html · ICML'19 ERL Workshop

add comment on HN