about
A Survey on Generative Diffusion Model (arxiv.org)
3 points by lapnect on Nov 5, 2024 | hide | past | pdf | discuss on HN

In plain words: Diffusion models turn random noise into images, text, speech, and more by cleaning it up step by step. A survey of the field lays out how they work, how they've been improved, and where they're used, plus open problems.

Abstract

Deep generative models have unlocked another profound realm of human creativity. By capturing and generalizing patterns within data, we have entered the epoch of all-encompassing Artificial Intelligence for General Creativity (AIGC). Notably, diffusion models, recognized as one of the paramount generative models, materialize human ideation into tangible instances across diverse domains, encompassing imagery, text, speech, biology, and healthcare. To provide advanced and comprehensive insights into diffusion, this survey comprehensively elucidates its developmental trajectory and future directions from three distinct angles: the fundamental formulation of diffusion, algorithmic enhancements, and the manifold applications of diffusion. Each layer is meticulously explored to offer a profound comprehension of its evolution. Structured and summarized approaches are presented in https://github.com/chq1155/A-Survey-on-Generative-Diffusion-Model.

Hanqun Cao, Cheng Tan, Zhangyang Gao, Yilun Xu, Guangyong Chen, Pheng-Ann Heng, Stan Z. Li
arXiv:2209.02646 · cs.AI · submitted Sep 6, 2022 · updated Dec 23, 2023
abstract · pdf · html

add comment on HN