about
Text to Photo-Realistic Image Synthesis with Generative Adversarial Networks (arxiv.org)
27 points by alando46 on Oct 20, 2017 | hide | past | pdf | 3 comments on HN

In plain words: Two networks work in sequence: one sketches rough shape and colors from a text description, the other turns that sketch into a 256x256 image with fine detail. Earlier one-pass systems lost detail; this two-step approach makes far more photo-realistic images.

Abstract · StackGAN: Text to Photo-realistic Image Synthesis with Stacked Generative Adversarial Networks

Synthesizing high-quality images from text descriptions is a challenging problem in computer vision and has many practical applications. Samples generated by existing text-to-image approaches can roughly reflect the meaning of the given descriptions, but they fail to contain necessary details and vivid object parts. In this paper, we propose Stacked Generative Adversarial Networks (StackGAN) to generate 256x256 photo-realistic images conditioned on text descriptions. We decompose the hard problem into more manageable sub-problems through a sketch-refinement process. The Stage-I GAN sketches the primitive shape and colors of the object based on the given text description, yielding Stage-I low-resolution images. The Stage-II GAN takes Stage-I results and text descriptions as inputs, and generates high-resolution images with photo-realistic details. It is able to rectify defects in Stage-I results and add compelling details with the refinement process. To improve the diversity of the synthesized images and stabilize the training of the conditional-GAN, we introduce a novel Conditioning Augmentation technique that encourages smoothness in the latent conditioning manifold. Extensive experiments and comparisons with state-of-the-arts on benchmark datasets demonstrate that the proposed method achieves significant improvements on generating photo-realistic images conditioned on text descriptions.

Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, Dimitris Metaxas
arXiv:1612.03242 · cs.CV, cs.AI, stat.ML · submitted Dec 10, 2016 · updated Aug 5, 2017
abstract · pdf · html · ICCV 2017 Oral Presentation

add comment on HN
Also discussed: Feb 2017 (2 points, 0 comments) · Jan 2017 (3 points, 0 comments) · Dec 2016 (2 points, 1 comment) · Dec 2016 (2 points, 0 comments) · Dec 2016 (3 points, 0 comments) · Dec 2016 (3 points, 1 comment)

Some interesting limitations here, but overall great work. The output is only 256x256, which doesn't seem useful in the short run. It would be interesting to see if there was another deep learning network that could fill in the pixels and net a higher resolution. It also looks like it uses models specifically trained for that given category (i.e. birds, flowers). I wonder how it would fair trying to visualize a distinct animal / flower that we only have a written description of.
The Github page shows a nice overview and some examples: https://github.com/hanzhanggit/StackGAN. Interestingly it seems to struggle with the birds' legs (e.g. merging them with branches, blurring them or just not generating them). Is this is a problem with the size of the feature or the colors or ...? I know very little about this field but I find it fascinating.
The Appendix of the pdf - https://arxiv.org/pdf/1612.03242 has more examples, including living rooms.