about
Deep Reinforcement Learning: An Overview (arxiv.org)
120 points by gwern on Feb 26, 2017 | hide | past | pdf | 6 comments on HN

In plain words: Deep reinforcement learning trains software to make decisions by trial and error, using deep neural networks to judge which actions pay off. This survey walks through its key parts, extra tricks like memory and transfer, and uses from games and robots to finance and healthcare.

Abstract

We give an overview of recent exciting achievements of deep reinforcement learning (RL). We discuss six core elements, six important mechanisms, and twelve applications. We start with background of machine learning, deep learning and reinforcement learning. Next we discuss core RL elements, including value function, in particular, Deep Q-Network (DQN), policy, reward, model, planning, and exploration. After that, we discuss important mechanisms for RL, including attention and memory, unsupervised learning, transfer learning, multi-agent RL, hierarchical RL, and learning to learn. Then we discuss various applications of RL, including games, in particular, AlphaGo, robotics, natural language processing, including dialogue systems, machine translation, and text generation, computer vision, neural architecture design, business management, finance, healthcare, Industry 4.0, smart grid, intelligent transportation systems, and computer systems. We mention topics not reviewed yet, and list a collection of RL resources. After presenting a brief summary, we close with discussions. Please see Deep Reinforcement Learning, arXiv:1810.06339, for a significant update.

Yuxi Li
arXiv:1701.07274 · cs.LG · submitted Jan 25, 2017 · updated Nov 26, 2018
abstract · pdf · html · Please see Deep Reinforcement Learning, arXiv:1810.06339, for a significant update

add comment on HN

What important DRL stuff does this not mention?
That DRL is incredibly difficult to stabilize in general.
More details, please. I'd especially be interested how the "in general" insight has been derived.
Basically, when deep reinforcement learning works, it's like magic, but unlike supervised learning where the default expectation is that it works without a hitch right out of the box, the default expectation for most new tasks through deep reinforcement learning is that it will fail, and you will need something to fix it.

For example, the high dimensionality of robotics makes it very difficult to apply deep reinforcement learning to it, although it definitely can and has been applied (and is, IMO, the future of robotics).

Another example is that simple supervised learning often outperforms DRL for many arcade games.

What is the current state of the art for multi-agent DRL?
>arxiv.org refused to connect.