about
High Efficiency Rl Agent (arxiv.org)
2 points by sel1 on Sep 2, 2019 | hide | past | pdf | discuss on HN

In plain words: The agent pairs a trial-and-error controller with a learned world model that remembers and predicts what happens next, so it can practice in imagination. It beat the usual approach on control tasks using fewer real interactions, and still worked when it couldn't see everything.

Abstract · Reinforcement learning with world model

Nowadays, model-free reinforcement learning algorithms have achieved remarkable performance on many decision making and control tasks, but high sample complexity and low sample efficiency still hinder the wide use of model-free reinforcement learning algorithms. In this paper, we argue that if we intend to design an intelligent agent that learns fast and transfers well, the agent must be able to reflect key elements of intelligence, like intuition, Memory, PredictionandCuriosity. We propose an agent framework that integrates off-policy reinforcement learning with world model learning, so as to embody the important features of intelligence in our algorithm design. We adopt the state-of-art model-free reinforcement learning algorithm, Soft Actor-Critic, as the agent intuition, and world model learning through RNN to endow the agent with memory, curiosity, and the ability to predict. We show that these ideas can work collaboratively with each other and our agent (RMC) can give new state-of-art results while maintaining sample efficiency and training stability. Moreover, our agent framework can be easily extended from MDP to POMDP problems without performance loss.

Jingbin Liu, Xinyang Gu, Shuai Liu
arXiv:1908.11494 · cs.AI · submitted Aug 30, 2019 · updated Oct 26, 2020
abstract · pdf · html

add comment on HN