In plain words: The agent acts as if the most promising untested possibility is true, and drops it when it fails. Given any limited set of possible worlds, it eventually acts as well as the best possible rule, and in simple certain cases gives error bounds.
Abstract · Optimistic Agents are Asymptotically Optimal
We use optimism to introduce generic asymptotically optimal reinforcement learning agents. They achieve, with an arbitrary finite or compact class of environments, asymptotically optimal behavior. Furthermore, in the finite deterministic case we provide finite error bounds.
Peter Sunehag, Marcus Hutter
arXiv:1210.0077 · cs.AI, cs.LG · submitted Sep 29, 2012
abstract · pdf · html · 13 LaTeX pages