about
Total stochastic gradient algorithms and applications in reinforcement learning (arxiv.org)
2 points by headalgorithm on Feb 6, 2019 | hide | past | pdf | discuss on HN

In plain words: Using the total derivative rule, this gives a way to build gradient estimates for graphs of variables, including one that stops at a middle step, not the score. It performs well in model-based reinforcement learning and offers clues to why one popular method is effective.

Abstract

Backpropagation and the chain rule of derivatives have been prominent; however, the total derivative rule has not enjoyed the same amount of attention. In this work we show how the total derivative rule leads to an intuitive visual framework for creating gradient estimators on graphical models. In particular, previous "policy gradient theorems" are easily derived. We derive new gradient estimators based on density estimation, as well as a likelihood ratio gradient, which "jumps" to an intermediate node, not directly to the objective function. We evaluate our methods on model-based policy gradient algorithms, achieve good performance, and present evidence towards demystifying the success of the popular PILCO algorithm.

Paavo Parmas
arXiv:1902.01722 · cs.LG, cs.AI, cs.NE, stat.ML · submitted Feb 5, 2019
abstract · pdf · html · NeurIPS 2018

add comment on HN