about
Is the Elephant Flying? Resolving Ambiguity in Text-to-Image Generative Models (arxiv.org)
1 point by PaulHoule on Nov 25, 2022 | hide | past | pdf | discuss on HN

In plain words: Picture generators were tested on tricky prompts where words could mean more than one thing, and a helper was built that asks the user a clarifying question before drawing. Its images matched what people actually meant more faithfully than drawing straight from the prompt.

Abstract · Is the Elephant Flying? Resolving Ambiguities in Text-to-Image Generative Models

Natural language often contains ambiguities that can lead to misinterpretation and miscommunication. While humans can handle ambiguities effectively by asking clarifying questions and/or relying on contextual cues and common-sense knowledge, resolving ambiguities can be notoriously hard for machines. In this work, we study ambiguities that arise in text-to-image generative models. We curate a benchmark dataset covering different types of ambiguities that occur in these systems. We then propose a framework to mitigate ambiguities in the prompts given to the systems by soliciting clarifications from the user. Through automatic and human evaluations, we show the effectiveness of our framework in generating more faithful images aligned with human intention in the presence of ambiguities.

Ninareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Varun Kumar, Qian Hu, Kai-Wei Chang, Richard Zemel, Aram Galstyan, Rahul Gupta
arXiv:2211.12503 · cs.CL, cs.CV, cs.LG, cs.MM · submitted Nov 17, 2022
abstract · pdf · html

add comment on HN