about
Observational Learning by Reinforcement Learning (arxiv.org)
60 points by guiambros on Jun 26, 2017 | hide | past | pdf | 2 comments on HN

In plain words: A learning agent watches another agent act in a shared world and picks up cues from how the world changes, without modeling that agent. In simple tasks its behavior changed after watching, and it used the teacher's actions when its reward depended on them.

Abstract

Observational learning is a type of learning that occurs as a function of observing, retaining and possibly replicating or imitating the behaviour of another agent. It is a core mechanism appearing in various instances of social learning and has been found to be employed in several intelligent species, including humans. In this paper, we investigate to what extent the explicit modelling of other agents is necessary to achieve observational learning through machine learning. Especially, we argue that observational learning can emerge from pure Reinforcement Learning (RL), potentially coupled with memory. Through simple scenarios, we demonstrate that an RL agent can leverage the information provided by the observations of an other agent performing a task in a shared environment. The other agent is only observed through the effect of its actions on the environment and never explicitly modeled. Two key aspects are borrowed from observational learning: i) the observer behaviour needs to change as a result of viewing a 'teacher' (another agent) and ii) the observer needs to be motivated somehow to engage in making use of the other agent's behaviour. The later is naturally modeled by RL, by correlating the learning agent's reward with the teacher agent's behaviour.

Diana Borsa, Bilal Piot, Rémi Munos, Olivier Pietquin
arXiv:1706.06617 · cs.LG, cs.AI, stat.ML · submitted Jun 20, 2017
abstract · pdf · html

add comment on HN

Copying behaviour without divining intent can lead to problems such as 'cargo-culting'.

Imagine observing a man shaking his leg, first one then the other, then his whole body convulses and twitches - is he dancing ?

Absent the knowledge that a wasp has flown up his trousers leg.

Copying without comprehension may lead to getting stung !

Inverse Reinforcement Learning [1] to reverse engineer goals will be needed especially for embodied AI in Partially Observed Enviroments, i.e. the real world (as opposed to simulations).

Berkeley's CS294-112 [2] Deep Reinforcement Learning for Robotics provides good coverage of methods of mirroring, DAGGer, Deep-Q, iLQR, and IRL.

[1] https://people.eecs.berkeley.edu/~pabbeel/cs287-fa12/slides/...

[2] https://www.youtube.com/playlist?list=PLkFD6_40KJIwTmSbCv9OV...

The man at risk of being stung is actively trying to prevent it. Blindly mimicking those movements, when not at risk of being stung, wouldn't have an effect, positively or negatively, on the ability to avoid stings. But if the man slammed his head into a post trying to avoid the wasp, copying would obviously be an issue. As long as the mimic is aware of harmful actions, learned either by experience (should be preferred) or given rules (cargo-culting risk), and stops copying the subject immediately (since it may have malicious intent or need emergency assistance), it should be safe, at least as safe as any of us can hope to be.