about
Discounted Reinforcement Learning Is Not an Optimization Problem (arxiv.org)
2 points by sel1 on Oct 9, 2019 | hide | past | pdf | discuss on HN

In plain words: When an agent with approximate value functions keeps discounting future rewards in never-ending tasks, no best policy exists, so the usual setup is not a real optimization problem. The analysis backs maximizing average reward per step instead.

Abstract

Discounted reinforcement learning is fundamentally incompatible with function approximation for control in continuing tasks. It is not an optimization problem in its usual formulation, so when using function approximation there is no optimal policy. We substantiate these claims, then go on to address some misconceptions about discounting and its connection to the average reward formulation. We encourage researchers to adopt rigorous optimization approaches, such as maximizing average reward, for reinforcement learning in continuing tasks.

Abhishek Naik, Roshan Shariff, Niko Yasui, Hengshuai Yao, Richard S. Sutton
arXiv:1910.02140 · cs.AI · submitted Oct 4, 2019 · updated Nov 27, 2019
abstract · pdf · html · Accepted for presentation at the Optimization Foundations of Reinforcement Learning Workshop at NeurIPS 2019

add comment on HN