about
DeepMind: Continuous control with deep reinforcement learning (arxiv.org)
1 point by vonnik on Sep 11, 2015 | hide | past | pdf | discuss on HN

In plain words: A trial-and-error system learns to pick smooth, continuous actions—like steering angles—rather than choosing from a fixed list, without being told the physics. With one set of settings it solved over 20 simulated physics tasks, matching a planner that knows the full rules.

Abstract · Continuous control with deep reinforcement learning

We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algorithm robustly solves more than 20 simulated physics tasks, including classic problems such as cartpole swing-up, dexterous manipulation, legged locomotion and car driving. Our algorithm is able to find policies whose performance is competitive with those found by a planning algorithm with full access to the dynamics of the domain and its derivatives. We further demonstrate that for many of the tasks the algorithm can learn policies end-to-end: directly from raw pixel inputs.

Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, Daan Wierstra
arXiv:1509.02971 · cs.LG, stat.ML · submitted Sep 9, 2015 · updated Jul 5, 2019
abstract · pdf · html · 10 pages + supplementary

add comment on HN
Also discussed: Sep 2015 (1 point, 0 comments)