about
Generating 512px photorealistic images and video with PixelCNNs (arxiv.org)
1 point by gwern on May 27, 2017 | hide | past | pdf | discuss on HN

In plain words: Instead of creating pixels one at a time, this generator fills in groups by treating some as independent once earlier ones are set. It samples in log N steps rather than N, making 512x512 images practical while staying close to the usual model's accuracy.

Abstract · Parallel Multiscale Autoregressive Density Estimation

PixelCNN achieves state-of-the-art results in density estimation for natural images. Although training is fast, inference is costly, requiring one network evaluation per pixel; O(N) for N pixels. This can be sped up by caching activations, but still involves generating each pixel sequentially. In this work, we propose a parallelized PixelCNN that allows more efficient inference by modeling certain pixel groups as conditionally independent. Our new PixelCNN model achieves competitive density estimation and orders of magnitude speedup - O(log N) sampling instead of O(N) - enabling the practical generation of 512x512 images. We evaluate the model on class-conditional image generation, text-to-image synthesis, and action-conditional video generation, showing that our model achieves the best results among non-pixel-autoregressive density models that allow efficient sampling.

Scott Reed, Aäron van den Oord, Nal Kalchbrenner, Sergio Gómez Colmenarejo, Ziyu Wang, Dan Belov, Nando de Freitas
arXiv:1703.03664 · cs.CV, cs.NE · submitted Mar 10, 2017
abstract · pdf · html

add comment on HN