In plain words: The agent starts out not knowing some parts of a decision problem even exist, then discovers them by exploring and getting help from an expert. It reaches near-optimal behavior, and remembering what it finds when new possibilities appear makes it learn faster.
Abstract
Methods for learning and planning in sequential decision problems often assume the learner is aware of all possible states and actions in advance. This assumption is sometimes untenable. In this paper, we give a method to learn factored markov decision problems from both domain exploration and expert assistance, which guarantees convergence to near-optimal behaviour, even when the agent begins unaware of factors critical to success. Our experiments show our agent learns optimal behaviour on small and large problems, and that conserving information on discovering new possibilities results in faster convergence.
Craig Innes, Alex Lascarides
arXiv:1902.10619 · cs.AI · submitted Feb 27, 2019
abstract · pdf · html