about
Generative Adversarial Text to Image Synthesis (arxiv.org)
1 point by EvgeniyZh on Aug 8, 2016 | hide | past | pdf | discuss on HN

In plain words: A system turns written descriptions into pictures using two competing networks: one draws images, the other judges how well they match the words. It produced believable birds and flowers from detailed descriptions, where earlier image makers handled only fixed categories like faces.

Abstract

Automatic synthesis of realistic images from text would be interesting and useful, but current AI systems are still far from this goal. However, in recent years generic and powerful recurrent neural network architectures have been developed to learn discriminative text feature representations. Meanwhile, deep convolutional generative adversarial networks (GANs) have begun to generate highly compelling images of specific categories, such as faces, album covers, and room interiors. In this work, we develop a novel deep architecture and GAN formulation to effectively bridge these advances in text and image model- ing, translating visual concepts from characters to pixels. We demonstrate the capability of our model to generate plausible images of birds and flowers from detailed text descriptions.

Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, Honglak Lee
arXiv:1605.05396 · cs.NE, cs.CV · submitted May 17, 2016 · updated Jun 5, 2016
abstract · pdf · html · ICML 2016

add comment on HN