about
Reinforcement Leaning in Feature Space: Matrix Bandit, Kernels, and Regret Bound (arxiv.org)
1 point by Anon84 on May 28, 2019 | hide | past | pdf | discuss on HN

In plain words: An agent learns a summary of how the world changes, then mixes trying actions with using what it knows to explore large spaces. Its loss grows with the square root of the number of steps, while earlier feature-based methods explode exponentially as tasks lengthen.

Abstract · Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret Bound

Exploration in reinforcement learning (RL) suffers from the curse of dimensionality when the state-action space is large. A common practice is to parameterize the high-dimensional value and policy functions using given features. However existing methods either have no theoretical guarantee or suffer a regret that is exponential in the planning horizon $H$. In this paper, we propose an online RL algorithm, namely the MatrixRL, that leverages ideas from linear bandit to learn a low-dimensional representation of the probability transition model while carefully balancing the exploitation-exploration tradeoff. We show that MatrixRL achieves a regret bound ${O}\big(H^2d\log T\sqrt{T}\big)$ where $d$ is the number of features. MatrixRL has an equivalent kernelized version, which is able to work with an arbitrary kernel Hilbert space without using explicit features. In this case, the kernelized MatrixRL satisfies a regret bound ${O}\big(H^2\widetilde{d}\log T\sqrt{T}\big)$, where $\widetilde{d}$ is the effective dimension of the kernel space. To our best knowledge, for RL using features or kernels, our results are the first regret bounds that are near-optimal in time $T$ and dimension $d$ (or $\widetilde{d}$) and polynomial in the planning horizon $H$.

Lin F. Yang, Mengdi Wang
arXiv:1905.10389 · cs.LG, stat.ML · submitted May 24, 2019 · updated Jun 13, 2019
abstract · pdf · html

add comment on HN