about
Diffusion Normalizing Flow (arxiv.org)
43 points by lnyan on Oct 15, 2021 | hide | past | pdf | 2 comments on HN

In plain words: A generator learns two linked noise processes: one turns data into random noise, the other turns noise back into data, trained together so they match. It captures sharper edges than normalizing flows and needs fewer sampling steps than standard diffusion models.

Abstract

We present a novel generative modeling method called diffusion normalizing flow based on stochastic differential equations (SDEs). The algorithm consists of two neural SDEs: a forward SDE that gradually adds noise to the data to transform the data into Gaussian random noise, and a backward SDE that gradually removes the noise to sample from the data distribution. By jointly training the two neural SDEs to minimize a common cost function that quantifies the difference between the two, the backward SDE converges to a diffusion process the starts with a Gaussian distribution and ends with the desired data distribution. Our method is closely related to normalizing flow and diffusion probabilistic models and can be viewed as a combination of the two. Compared with normalizing flow, diffusion normalizing flow is able to learn distributions with sharp boundaries. Compared with diffusion probabilistic models, diffusion normalizing flow requires fewer discretization steps and thus has better sampling efficiency. Our algorithm demonstrates competitive performance in both high-dimension data density estimation and image generation tasks.

Qinsheng Zhang, Yongxin Chen
arXiv:2110.07579 · cs.LG · submitted Oct 14, 2021
abstract · pdf · html · Neurips 2021

add comment on HN

Need to read this one, if you are the authors can I ask you to delineate it against Song et al. "Score-Based Generative Modeling through Stochastic Differential Equations" https://arxiv.org/abs/2011.13456 ?
Pretty cool in fact all the methods they demo there show good approximation of difficult distributions (if that looks easy, take a look at scikit learn’s manifold doc page). The shoe that hasn’t dropped in the article is behavior in high dimensions. For instance, I seem to recall that backwards integration of high dimensional DEs is unstable (not to mention memory issues).