In plain words: The system builds a map of the game's world, then uses it to shrink the list of commands and pick one by treating the choice as a question-answering task it can pre-train on. It learned to play faster than the standard alternatives.
Abstract
Text-based adventure games provide a platform on which to explore reinforcement learning in the context of a combinatorial action space, such as natural language. We present a deep reinforcement learning architecture that represents the game state as a knowledge graph which is learned during exploration. This graph is used to prune the action space, enabling more efficient exploration. The question of which action to take can be reduced to a question-answering task, a form of transfer learning that pre-trains certain parts of our architecture. In experiments using the TextWorld framework, we show that our proposed technique can learn a control policy faster than baseline alternatives. We have also open-sourced our code at https://github.com/rajammanabrolu/KG-DQN.
Prithviraj Ammanabrolu, Mark O. Riedl
arXiv:1812.01628 · cs.CL, cs.AI, cs.LG · submitted Dec 4, 2018 · updated Mar 25, 2019
abstract · pdf · html · Proceedings of NAACL-HLT 2019