In plain words: A human teacher gives a robot free-form commands while it learns from rewards and simple yes/no feedback, and the robot works out what each command means as it goes. Tests with a real robot show this speeds learning and needs fewer teaching signals.
Abstract
In this paper, we propose a framework that enables a human teacher to shape a robot behaviour by interactively providing it with unlabeled instructions. We ground the meaning of instruction signals in the task-learning process, and use them simultaneously for guiding the latter. We implement our framework as a modular architecture, named TICS (Task-Instruction-Contingency-Shaping) that combines different information sources: a predefined reward function, human evaluative feedback and unlabeled instructions. This approach provides a novel perspective for robotic task learning that lies between Reinforcement Learning and Supervised Learning paradigms. We evaluate our framework both in simulation and with a real robot. The experimental results demonstrate the effectiveness of our framework in accelerating the task-learning process and in reducing the number of required teaching signals.
Anis Najar, Olivier Sigaud, Mohamed Chetouani
arXiv:1902.01670 · cs.LG, cs.RO, stat.ML · submitted Feb 5, 2019 · updated Nov 24, 2020
abstract · pdf · html