about
Transformers are sample-efficient world models (arxiv.org)
67 points by lucidrains on Sep 2, 2022 | hide | past | pdf | 10 comments on HN

In plain words: The agent turns game images into tokens and uses a Transformer to predict what happens next, letting it practice in imagination. With two hours of Atari play, it beat humans on 10 of 26 games, ahead of earlier agents that don't plan ahead.

Abstract · Transformers are Sample-Efficient World Models

Deep reinforcement learning agents are notoriously sample inefficient, which considerably limits their application to real-world problems. Recently, many model-based methods have been designed to address this issue, with learning in the imagination of a world model being one of the most prominent approaches. However, while virtually unlimited interaction with a simulated environment sounds appealing, the world model has to be accurate over extended periods of time. Motivated by the success of Transformers in sequence modeling tasks, we introduce IRIS, a data-efficient agent that learns in a world model composed of a discrete autoencoder and an autoregressive Transformer. With the equivalent of only two hours of gameplay in the Atari 100k benchmark, IRIS achieves a mean human normalized score of 1.046, and outperforms humans on 10 out of 26 games, setting a new state of the art for methods without lookahead search. To foster future research on Transformers and world models for sample-efficient reinforcement learning, we release our code and models at https://github.com/eloialonso/iris.

Vincent Micheli, Eloi Alonso, François Fleuret
arXiv:2209.00588 · cs.LG, cs.AI, cs.CV · submitted Sep 1, 2022 · updated Mar 1, 2023
abstract · pdf · html · ICLR 2023 (notable top 5%)

add comment on HN
Also discussed: Sep 2022 (7 points, 0 comments)

>With the equivalent of only two hours of gameplay in the Atari 100k benchmark

Does that include Montezuma's Revenge? Because that's the Atari game RL agents completely fail at.

Edit: Yup, Atari 100k doesn't include MR.

https://twitter.com/arankomatsuzaki/status/14553550319019089...

Hmm. So what they show is that Transformers work exceptionally well for some games, but not in general.

If you allow MCTS, then EfficientZero performs +85% better. And if you look at the median (instead of the mean), then SPR performs +37% better.

But for some games, like ChopperCommand or Krull, IRIS (this paper) performs beyond comparison to other methods.

You wonder if that means intelligence really works like other bodily systems. Animals work so well not because of a single trick, but due to a big bag of tricks. Not "one single gene" or algorithms learning good performance at everything, but 1000 different algorithms and a good way of selecting between them. Not 1 or 2 or 10, but a large number.
I guess androids do dream of electric sheep.
Can anyone dumb this down? Also is it a big deal?
Reinforcement learning has an issue currently with modeling complicated aspects of an agent’s given simulation (their “world model”) over long time horizons.

This is probably because there is a vast amount of information in any given “frame” of this simulation and the agent has no baked-in priors about its environment (or if they do, it’s because they’ve been hand-programmed, which isn’t ideal in a field premised on the automation of such things).

If you instead pretrain _another_ model using the same tech as GPT3 on the agent’s environment, the agents can use this model during their simulations to model the world they see, essentially giving them something that looks kind of like common sense (although I certainly wouldn’t call it that to another researcher).

That title could do with a hyphen.
Ok, let's have one.
You rock dang!
Here you go: -