In plain words: A speaker describes images to listeners with different ideas and senses, so it must figure out what each one can grasp and adjust to new partners. Guessing what listeners understand made it win more often than speaking the same way to everyone.
Abstract
An agent who interacts with a wide population of other agents needs to be aware that there may be variations in their understanding of the world. Furthermore, the machinery which they use to perceive may be inherently different, as is the case between humans and machines. In this work, we present both an image reference game between a speaker and a population of listeners where reasoning about the concepts other agents can comprehend is necessary and a model formulation with this capability. We focus on reasoning about the conceptual understanding of others, as well as adapting to novel gameplay partners and dealing with differences in perceptual machinery. Our experiments on three benchmark image/attribute datasets suggest that our learner indeed encodes information directly pertaining to the understanding of other agents, and that leveraging this information is crucial for maximizing gameplay performance.
Rodolfo Corona, Stephan Alaniz, Zeynep Akata
arXiv:1910.04872 · cs.AI · submitted Oct 10, 2019 · updated Nov 19, 2019
abstract · pdf · html · Published in NeurIPS 2019