In plain words: A 4-billion-part network chops images into patches and treats them like words, so one system learns how text and pictures match and can draw new pictures from a caption. It beat older noise-to-image models and DALL-E on the standard caption-to-picture quality test.
Abstract · CogView: Mastering Text-to-Image Generation via Transformers
Text-to-Image generation in the general domain has long been an open problem, which requires both a powerful generative model and cross-modal understanding. We propose CogView, a 4-billion-parameter Transformer with VQ-VAE tokenizer to advance this problem. We also demonstrate the finetuning strategies for various downstream tasks, e.g. style learning, super-resolution, text-image ranking and fashion design, and methods to stabilize pretraining, e.g. eliminating NaN losses. CogView achieves the state-of-the-art FID on the blurred MS COCO dataset, outperforming previous GAN-based models and a recent similar work DALL-E.
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, Jie Tang
arXiv:2105.13290 · cs.CV, cs.LG · submitted May 26, 2021 · updated Nov 5, 2021
abstract · pdf · html · to appear in NeurIPS 2021