about
Making Efficient Use of Demonstrations to Solve Hard Exploration Problems (arxiv.org)
2 points by sel1 on Sep 5, 2019 | hide | past | pdf | discuss on HN

In plain words: A computer player uses a few human demonstrations to guide its exploration in games where it can't see everything and each run starts differently. It solved several of eight new tasks where other players never found one success in tens of billions of tries.

Abstract

This paper introduces R2D3, an agent that makes efficient use of demonstrations to solve hard exploration problems in partially observable environments with highly variable initial conditions. We also introduce a suite of eight tasks that combine these three properties, and show that R2D3 can solve several of the tasks where other state of the art methods (both with and without demonstrations) fail to see even a single successful trajectory after tens of billions of steps of exploration.

Tom Le Paine, Caglar Gulcehre, Bobak Shahriari, Misha Denil, Matt Hoffman, Hubert Soyer, Richard Tanburn, Steven Kapturowski, Neil Rabinowitz, Duncan Williams, Gabriel Barth-Maron, Ziyu Wang, et al.
arXiv:1909.01387 · cs.LG, cs.AI · submitted Sep 3, 2019
abstract · pdf · html

add comment on HN