In plain words: Image generators usually learn by slowly undoing random noise added to pictures. The same recipe works with plain, non-random damage like blur or masking, still producing good images and suggesting noise is not what makes these generators work.
Abstract
Standard diffusion models involve an image transform -- adding Gaussian noise -- and an image restoration operator that inverts this degradation. We observe that the generative behavior of diffusion models is not strongly dependent on the choice of image degradation, and in fact an entire family of generative models can be constructed by varying this choice. Even when using completely deterministic degradations (e.g., blur, masking, and more), the training and test-time update rules that underlie diffusion models can be easily generalized to create generative models. The success of these fully deterministic models calls into question the community's understanding of diffusion models, which relies on noise in either gradient Langevin dynamics or variational inference, and paves the way for generalized diffusion models that invert arbitrary processes. Our code is available at https://github.com/arpitbansal297/Cold-Diffusion-Models
Arpit Bansal, Eitan Borgnia, Hong-Min Chu, Jie S. Li, Hamid Kazemi, Furong Huang, Micah Goldblum, Jonas Geiping, Tom Goldstein
arXiv:2208.09392 · cs.CV, cs.LG · submitted Aug 19, 2022
abstract · pdf · html
I wonder to what degree certain parts of diffusion dictate using certain noise functions, and how much this paper truly challenges how we understand them. Cool to see it was researched.
Next idea: it seems like a lot of steps could be skipped by using things like momentum during the inference time. I'm sure OpenAI has already implemented several clever tricks like that in production for DallE.
I'm working on (various, non-diffusion) methods for 2D drawing to 3D output right now.