about
Playing hard exploration games by watching YouTube (arxiv.org)
127 points by indescions_2018 on May 30, 2018 | hide | past | pdf | 11 comments on HN

In plain words: It learns a shared picture-and-sound representation from messy, unaligned gameplay videos, then turns one YouTube clip into a goal signal that rewards copying. Unlike usual imitation, which needs the player's actions and rewards, it beat human scores on three hard games without game rewards.

Abstract

Deep reinforcement learning methods traditionally struggle with tasks where environment rewards are particularly sparse. One successful method of guiding exploration in these domains is to imitate trajectories provided by a human demonstrator. However, these demonstrations are typically collected under artificial conditions, i.e. with access to the agent's exact environment setup and the demonstrator's action and reward trajectories. Here we propose a two-stage method that overcomes these limitations by relying on noisy, unaligned footage without access to such data. First, we learn to map unaligned videos from multiple sources to a common representation using self-supervised objectives constructed over both time and modality (i.e. vision and sound). Second, we embed a single YouTube video in this representation to construct a reward function that encourages an agent to imitate human gameplay. This method of one-shot imitation allows our agent to convincingly exceed human-level performance on the infamously hard exploration games Montezuma's Revenge, Pitfall! and Private Eye for the first time, even if the agent is not presented with any environment rewards.

Yusuf Aytar, Tobias Pfaff, David Budden, Tom Le Paine, Ziyu Wang, Nando de Freitas
arXiv:1805.11592 · cs.LG, cs.AI, cs.CV, stat.ML · submitted May 29, 2018 · updated Nov 30, 2018
abstract · pdf · html

add comment on HN

Neat, I don't understand what they mean by having embedded a reward video into the set. Is that a video where copying the behaviour will deliver victory?
Yes, they take the state of the video every 16 frames and look at its embedding. These were made into checkpoints.

The AI is rewarded if at each checkpoint the state vector its produced is sufficiently aligned with the videos.

I guess that's the initial training to deal with sparse rewards.

here's video of the agent actually playing (linked in the paper): https://www.youtube.com/watch?v=Msy82sIfprI
This is really cool. A step in the right direction towards general learning through observation.
This is actually quite human. I also watch Let's plays if I struggle with a quest (or game in general).

Also interesting assumption to say "harder = fewer rewards". Probably doesn't always apply but is a good generalization.

Are audio cues also analyzed here? ie: "We observe that use of the audio signal in CMC results in more emphasis being placed on key items and their location in the inventory"
Yes. s3.2 suggests they use audio cues to help them align the video frames from different videos. (I guess it's easier to correlate audio than video.)
I can imagine video quality on youtube varies more than audio, or that audio is easier to hash / make signatures of.
Hmm I was more under the impression it was for context, in other words creating/executing strategies based on what audio cues were received (like when a key or coin is acquired) and keeping tabs of what actions the user performed after that point.
This should probably say "ML" or "AI" or whatever, I was slightly disappointed to realize it was not a funny paper about… I don't know to be fair.
I can see where you're coming from, the title definitely made me initially feel like it was going to be about getting satisfaction from watching other people play a game or something like that.