about
Tutorial on Diffusion Models for Imaging and Vision (arxiv.org)
2 points by qwertyforce on Apr 2, 2024 | hide | past | pdf | discuss on HN

In plain words: A teaching guide explains how diffusion models work: they learn to turn random noise step by step into images or videos, the trick behind today's text-to-image tools. It is written for students who want to build or use these models.

Abstract

The astonishing growth of generative tools in recent years has empowered many exciting applications in text-to-image generation and text-to-video generation. The underlying principle behind these generative tools is the concept of diffusion, a particular sampling mechanism that has overcome some shortcomings that were deemed difficult in the previous approaches. The goal of this tutorial is to discuss the essential ideas underlying the diffusion models. The target audience of this tutorial includes undergraduate and graduate students who are interested in doing research on diffusion models or applying these models to solve other problems.

Stanley H. Chan
arXiv:2403.18103 · cs.LG, cs.CV · submitted Mar 26, 2024 · updated Jan 8, 2025
abstract · pdf · html

add comment on HN
Also discussed: Sep 2024 (221 points, 18 comments)