about
Evolving Reinforcement Learning Algorithms (arxiv.org)
8 points by jonbaer on Jan 18, 2021 | hide | past | pdf | discuss on HN

In plain words: A computer searches formulas for how an agent should learn from rewards. From scratch it reinvented temporal-difference learning (learning from each step's own prediction); starting from a standard game agent, it found tweaks that generalized better than the usual fixed rule to new games.

Abstract

We propose a method for meta-learning reinforcement learning algorithms by searching over the space of computational graphs which compute the loss function for a value-based model-free RL agent to optimize. The learned algorithms are domain-agnostic and can generalize to new environments not seen during training. Our method can both learn from scratch and bootstrap off known existing algorithms, like DQN, enabling interpretable modifications which improve performance. Learning from scratch on simple classical control and gridworld tasks, our method rediscovers the temporal-difference (TD) algorithm. Bootstrapped from DQN, we highlight two learned algorithms which obtain good generalization performance over other classical control tasks, gridworld type tasks, and Atari games. The analysis of the learned algorithm behavior shows resemblance to recently proposed RL algorithms that address overestimation in value-based methods.

John D. Co-Reyes, Yingjie Miao, Daiyi Peng, Esteban Real, Sergey Levine, Quoc V. Le, Honglak Lee, Aleksandra Faust
arXiv:2101.03958 · cs.LG, cs.AI, cs.NE · submitted Jan 8, 2021 · updated Nov 10, 2022
abstract · pdf · html · ICLR 2021 Oral. See project website at https://sites.google.com/view/evolvingrl

add comment on HN
Also discussed: Feb 2021 (1 point, 0 comments) · Jan 2021 (2 points, 0 comments)