about
Hierarchical Decision Making by Generating and Following Natural Language (arxiv.org)
3 points by jonbaer on Sep 16, 2019 | hide | past | pdf | discuss on HN

In plain words: An agent first writes a plan in plain language, then a second model carries it out to coordinate many units over long stretches of a strategy game. Trained on 76,000 human instruction-and-play pairs, it beat copying human moves directly, and language's word-combining structure mattered.

Abstract · Hierarchical Decision Making by Generating and Following Natural Language Instructions

We explore using latent natural language instructions as an expressive and compositional representation of complex actions for hierarchical decision making. Rather than directly selecting micro-actions, our agent first generates a latent plan in natural language, which is then executed by a separate model. We introduce a challenging real-time strategy game environment in which the actions of a large number of units must be coordinated across long time scales. We gather a dataset of 76 thousand pairs of instructions and executions from human play, and train instructor and executor models. Experiments show that models using natural language as a latent variable significantly outperform models that directly imitate human actions. The compositional structure of language proves crucial to its effectiveness for action representation. We also release our code, models and data.

Hengyuan Hu, Denis Yarats, Qucheng Gong, Yuandong Tian, Mike Lewis
arXiv:1906.00744 · cs.AI, cs.CL · submitted Jun 3, 2019 · updated Oct 2, 2019
abstract · pdf · html

add comment on HN