In plain words: Diffusion models make images by starting with random noise and cleaning it up step by step; this paper improves the network design and uses a classifier's feedback to steer each step toward more realistic pictures. On ImageNet they beat the previous best generator, GANs, in quality while covering more variety, matching BigGAN-deep in as few as 25 cleanup steps.
Abstract
We show that diffusion models can achieve image sample quality superior to the current state-of-the-art generative models. We achieve this on unconditional image synthesis by finding a better architecture through a series of ablations. For conditional image synthesis, we further improve sample quality with classifier guidance: a simple, compute-efficient method for trading off diversity for fidelity using gradients from a classifier. We achieve an FID of 2.97 on ImageNet 128$\times$128, 4.59 on ImageNet 256$\times$256, and 7.72 on ImageNet 512$\times$512, and we match BigGAN-deep even with as few as 25 forward passes per sample, all while maintaining better coverage of the distribution. Finally, we find that classifier guidance combines well with upsampling diffusion models, further improving FID to 3.94 on ImageNet 256$\times$256 and 3.85 on ImageNet 512$\times$512. We release our code at https://github.com/openai/guided-diffusion
Prafulla Dhariwal, Alex Nichol
arXiv:2105.05233 · cs.LG, cs.AI, cs.CV, stat.ML · submitted May 11, 2021 · updated Jun 1, 2021
abstract · pdf · html · Added compute requirements, ImageNet 256$\times$256 upsampling FID and samples, DDIM guided sampler, fixed typos
With diffusion models, you need to do >25 forward passes to achieve a result. It’s kind of like an O(1) algorithm vs O(N): stylegan has one pass, diffusion models have N. And N is currently 25 or more, which means it tends to be 25x slower than stylegan at a minimum. (In our experience it was often many seconds to a full minute before we saw results, but we didn’t try very hard to make it fast, and this paper shows advances in speed since then.)
The flipside: this paper has the most beautiful photorealistic complex samples I’ve ever seen. I don’t really care if they’re cherry picked; it’s hard to pick cherries on a rotten tree. I know.
The last thing I want to say is that I am disappointed the model is, yet again, not released. The hoarding of ML models needs to stop. These research models aren’t commercially useful, but they have incredible benefits for people outside of ML. I myself got into ML thanks to being able to play with GPT-2 and Sfylegan. I would leap at the chance to play with AlphaZero or OpenAI’s old Dota 2 models; no reason to keep those locked up. I also don’t care about the excuses for not releasing: none of them hold water. In my experience, no model has ever had a society-threatening impact, and it’s usually a euphemism for “we’re worried there’s a small chance we might look bad.” Releasing your model is as simple as scp’ing to your server; just do it and quit worrying so much.
Love the work. Diffusion models are a really interesting way of thinking about generative approaches in ML.