about
Visual Semantic Planning Using Deep Successor Representations (arxiv.org)
2 points by gwern on Feb 22, 2018 | hide | past | pdf | 1 comment on HN

In plain words: An agent learns to plan action sequences that turn one scene into a goal scene, by copying demonstrations and practicing on its own, with a model that predicts how scenes lead to later scenes. It reached near-optimal results across many tasks in a simulated home.

Abstract · Visual Semantic Planning using Deep Successor Representations

A crucial capability of real-world intelligent agents is their ability to plan a sequence of actions to achieve their goals in the visual world. In this work, we address the problem of visual semantic planning: the task of predicting a sequence of actions from visual observations that transform a dynamic environment from an initial state to a goal state. Doing so entails knowledge about objects and their affordances, as well as actions and their preconditions and effects. We propose learning these through interacting with a visual and dynamic environment. Our proposed solution involves bootstrapping reinforcement learning with imitation learning. To ensure cross task generalization, we develop a deep predictive model based on successor representations. Our experimental results show near optimal results across a wide range of tasks in the challenging THOR environment.

Yuke Zhu, Daniel Gordon, Eric Kolve, Dieter Fox, Li Fei-Fei, Abhinav Gupta, Roozbeh Mottaghi, Ali Farhadi
arXiv:1705.08080 · cs.CV, cs.LG, cs.RO · submitted May 23, 2017 · updated Aug 15, 2017
abstract · pdf · html · ICCV 2017 camera ready

add comment on HN