In plain words: A learning agent watches another agent act in a shared world and picks up cues from how the world changes, without modeling that agent. In simple tasks its behavior changed after watching, and it used the teacher's actions when its reward depended on them.
Abstract
Observational learning is a type of learning that occurs as a function of observing, retaining and possibly replicating or imitating the behaviour of another agent. It is a core mechanism appearing in various instances of social learning and has been found to be employed in several intelligent species, including humans. In this paper, we investigate to what extent the explicit modelling of other agents is necessary to achieve observational learning through machine learning. Especially, we argue that observational learning can emerge from pure Reinforcement Learning (RL), potentially coupled with memory. Through simple scenarios, we demonstrate that an RL agent can leverage the information provided by the observations of an other agent performing a task in a shared environment. The other agent is only observed through the effect of its actions on the environment and never explicitly modeled. Two key aspects are borrowed from observational learning: i) the observer behaviour needs to change as a result of viewing a 'teacher' (another agent) and ii) the observer needs to be motivated somehow to engage in making use of the other agent's behaviour. The later is naturally modeled by RL, by correlating the learning agent's reward with the teacher agent's behaviour.
Diana Borsa, Bilal Piot, Rémi Munos, Olivier Pietquin
arXiv:1706.06617 · cs.LG, cs.AI, stat.ML · submitted Jun 20, 2017
abstract · pdf · html
Imagine observing a man shaking his leg, first one then the other, then his whole body convulses and twitches - is he dancing ?
Absent the knowledge that a wasp has flown up his trousers leg.
Copying without comprehension may lead to getting stung !
Inverse Reinforcement Learning [1] to reverse engineer goals will be needed especially for embodied AI in Partially Observed Enviroments, i.e. the real world (as opposed to simulations).
Berkeley's CS294-112 [2] Deep Reinforcement Learning for Robotics provides good coverage of methods of mirroring, DAGGer, Deep-Q, iLQR, and IRL.
[1] https://people.eecs.berkeley.edu/~pabbeel/cs287-fa12/slides/...
[2] https://www.youtube.com/playlist?list=PLkFD6_40KJIwTmSbCv9OV...