In plain words: To help robots learn control with fewer real trials, past experiences are reused in two new ways: mirrored when the task is symmetric, and relabeled with looser goals. Early tests show learning speeds up a lot compared with training on experiences as collected.
Abstract · Towards More Sample Efficiency in Reinforcement Learning with Data Augmentation
Deep reinforcement learning (DRL) is a promising approach for adaptive robot control, but its current application to robotics is currently hindered by high sample requirements. We propose two novel data augmentation techniques for DRL in order to reuse more efficiently observed data. The first one called Kaleidoscope Experience Replay exploits reflectional symmetries, while the second called Goal-augmented Experience Replay takes advantage of lax goal definitions. Our preliminary experimental results show a large increase in learning speed.
Yijiong Lin, Jiancong Huang, Matthieu Zimmer, Juan Rojas, Paul Weng
arXiv:1910.09959 · cs.AI, cs.RO · submitted Oct 19, 2019 · updated Nov 15, 2019
abstract · pdf · html · NeurIPS 2019 Workshop on Robot Learning: Control and Interaction in the Real World (accepted after double-blind peer review). arXiv admin note: substantial text overlap with arXiv:1909.10707