about
Noisy Networks for Exploration (arxiv.org)
3 points by guiambros on Jul 9, 2017 | hide | past | pdf | discuss on HN

In plain words: Random noise is added to the network's weights and tuned during learning, so the agent explores on its own instead of using hand-set tricks like random action picks. This scored higher across many Atari games, sometimes moving from below to above human level.

Abstract

We introduce NoisyNet, a deep reinforcement learning agent with parametric noise added to its weights, and show that the induced stochasticity of the agent's policy can be used to aid efficient exploration. The parameters of the noise are learned with gradient descent along with the remaining network weights. NoisyNet is straightforward to implement and adds little computational overhead. We find that replacing the conventional exploration heuristics for A3C, DQN and dueling agents (entropy reward and $ε$-greedy respectively) with NoisyNet yields substantially higher scores for a wide range of Atari games, in some cases advancing the agent from sub to super-human performance.

Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, Shane Legg
arXiv:1706.10295 · cs.LG, stat.ML · submitted Jun 30, 2017 · updated Jul 9, 2019
abstract · pdf · html · ICLR 2018

add comment on HN
Also discussed: Jul 2019 (1 point, 0 comments) · Jul 2017 (6 points, 0 comments)