about
A0C: Alpha Zero in Continuous Action Space (arxiv.org)
6 points by jonbaer on May 27, 2018 | hide | past | pdf | discuss on HN

In plain words: AlphaZero usually picks moves from a fixed list, so this study reworks its blend of tree search and neural networks to handle actions that are any number, like a robot's steering angle. Early tests on swinging up a pendulum show the approach can work.

Abstract

A core novelty of Alpha Zero is the interleaving of tree search and deep learning, which has proven very successful in board games like Chess, Shogi and Go. These games have a discrete action space. However, many real-world reinforcement learning domains have continuous action spaces, for example in robotic control, navigation and self-driving cars. This paper presents the necessary theoretical extensions of Alpha Zero to deal with continuous action space. We also provide some preliminary experiments on the Pendulum swing-up task, empirically showing the feasibility of our approach. Thereby, this work provides a first step towards the application of iterated search and learning in domains with a continuous action space.

Thomas M. Moerland, Joost Broekens, Aske Plaat, Catholijn M. Jonker
arXiv:1805.09613 · stat.ML, cs.AI, cs.LG, cs.RO, eess.SY · submitted May 24, 2018
abstract · pdf · html

add comment on HN