In plain words: Agents play hide-and-seek among obstacles, seeing only a first-person view, and learn from scratch to evade capture, apparently by judging when they can be seen. Weakening them, such as by slowing them down, makes training harder but yields more useful learned features.
Abstract
We train embodied agents to play Visual Hide and Seek where a prey must navigate in a simulated environment in order to avoid capture from a predator. We place a variety of obstacles in the environment for the prey to hide behind, and we only give the agents partial observations of their environment using an egocentric perspective. Although we train the model to play this game from scratch, experiments and visualizations suggest that the agent learns to predict its own visibility in the environment. Furthermore, we quantitatively analyze how agent weaknesses, such as slower speed, effect the learned policy. Our results suggest that, although agent weaknesses make the learning problem more challenging, they also cause more useful features to be learned. Our project website is available at: http://www.cs.columbia.edu/ ~bchen/visualhideseek/.
Boyuan Chen, Shuran Song, Hod Lipson, Carl Vondrick
arXiv:1910.07882 · cs.AI, cs.CV, cs.LG, cs.MA, cs.RO · submitted Oct 15, 2019
abstract · pdf · html · 14 pages, 8 figures