about
Task-Relevant Adversarial Imitation Learning (arxiv.org)
3 points by sel1 on Oct 4, 2019 | hide | past | pdf | discuss on HN

In plain words: Robots learning from human demonstrations can have their judge network fix on meaningless visual clues instead of what matters, giving useless feedback. Constraining the judge to focus on task-relevant details let robots master arm-and-hand tasks from camera images, beating plain copying and adversarial imitation.

Abstract

We show that a critical vulnerability in adversarial imitation is the tendency of discriminator networks to learn spurious associations between visual features and expert labels. When the discriminator focuses on task-irrelevant features, it does not provide an informative reward signal, leading to poor task performance. We analyze this problem in detail and propose a solution that outperforms standard Generative Adversarial Imitation Learning (GAIL). Our proposed method, Task-Relevant Adversarial Imitation Learning (TRAIL), uses constrained discriminator optimization to learn informative rewards. In comprehensive experiments, we show that TRAIL can solve challenging robotic manipulation tasks from pixels by imitating human operators without access to any task rewards, and clearly outperforms comparable baseline imitation agents, including those trained via behaviour cloning and conventional GAIL.

Konrad Zolna, Scott Reed, Alexander Novikov, Sergio Gomez Colmenarejo, David Budden, Serkan Cabi, Misha Denil, Nando de Freitas, Ziyu Wang
arXiv:1910.01077 · cs.LG, cs.AI, cs.RO, stat.ML · submitted Oct 2, 2019 · updated Nov 12, 2020
abstract · pdf · html · Accepted to CoRL 2020 (see presentation here: https://youtu.be/ZgQvFGuEgFU )

add comment on HN