about
Striving for Simplicity in Off-Policy Deep Reinforcement Learning (arxiv.org)
36 points by jonbaer on Jul 12, 2019 | hide | past | pdf | 2 comments on HN

In plain words: Instead of letting a game agent practice live, they train it only on a fixed log of one agent's past play across 60 Atari games. Their new trick, which averages several guesses of future rewards, beats both the original agent and other offline learners.

Abstract · An Optimistic Perspective on Offline Reinforcement Learning

Off-policy reinforcement learning (RL) using a fixed offline dataset of logged interactions is an important consideration in real world applications. This paper studies offline RL using the DQN replay dataset comprising the entire replay experience of a DQN agent on 60 Atari 2600 games. We demonstrate that recent off-policy deep RL algorithms, even when trained solely on this fixed dataset, outperform the fully trained DQN agent. To enhance generalization in the offline setting, we present Random Ensemble Mixture (REM), a robust Q-learning algorithm that enforces optimal Bellman consistency on random convex combinations of multiple Q-value estimates. Offline REM trained on the DQN replay dataset surpasses strong RL baselines. Ablation studies highlight the role of offline dataset size and diversity as well as the algorithm choice in our positive results. Overall, the results here present an optimistic view that robust RL algorithms trained on sufficiently large and diverse offline datasets can lead to high quality policies. The DQN replay dataset can serve as an offline RL benchmark and is open-sourced.

Rishabh Agarwal, Dale Schuurmans, Mohammad Norouzi
arXiv:1907.04543 · cs.LG, cs.AI, stat.ML · submitted Jul 10, 2019 · updated Jun 22, 2020
abstract · pdf · html · ICML 2020. An earlier version was titled "Striving for Simplicity in Off-Policy Deep Reinforcement Learning". Project Website: https://offline-rl.github.io

add comment on HN

I wrote a quick summary on /r/reinforcementlearning: https://www.reddit.com/r/reinforcementlearning/comments/cc9g...
This is a dense read, and I will have to spend a lot of time on it, but is it really surprising that Atari can be learned from 50m events? This isn't a real domain, and it's not a complex domain and it's not a domain that synthesises the challenges of real domains (for example noise between the control input and outputs, noise in the sensors, dynamism in the domain)

That's ok - but let's not conflate "learning Atari" and "reinforcement learning"