about
Tutorial on diffusion models for imaging and vision (arxiv.org)
221 points by Anon84 on Sep 10, 2024 | hide | past | pdf | 18 comments on HN

In plain words: A teaching guide explains how diffusion models work: they learn to turn random noise step by step into images or videos, the trick behind today's text-to-image tools. It is written for students who want to build or use these models.

Abstract · Tutorial on Diffusion Models for Imaging and Vision

The astonishing growth of generative tools in recent years has empowered many exciting applications in text-to-image generation and text-to-video generation. The underlying principle behind these generative tools is the concept of diffusion, a particular sampling mechanism that has overcome some shortcomings that were deemed difficult in the previous approaches. The goal of this tutorial is to discuss the essential ideas underlying the diffusion models. The target audience of this tutorial includes undergraduate and graduate students who are interested in doing research on diffusion models or applying these models to solve other problems.

Stanley H. Chan
arXiv:2403.18103 · cs.LG, cs.CV · submitted Mar 26, 2024 · updated Jan 8, 2025
abstract · pdf · html

add comment on HN
Also discussed: Apr 2024 (2 points, 0 comments)

A very useful guide about how diffusion models work and implementation: https://keras.io/examples/generative/ddim/

I find the explanation in this article very intuitive.

So anyone up for a discussion on this paper at: https://www.alphaxiv.org/abs/2403.18103
Does something similar exist for LLM/GPT?

Edit to add: I'm mostly interested in this aspect:

"The target audience of this tutorial includes [those] who are interested in [...] applying these models to solve other problems."

Andrej Karpathy has a youtube playlist:

https://www.youtube.com/playlist?list=PLAqhIrjkxbuWI23v9cThs...

He is building new learning materials under his new company "Eureka Labs":

https://eurekalabs.ai

Sebastian Raschka's book "Build a Large Language Model (From Scratch) just released:

https://www.manning.com/books/build-a-large-language-model-f...

All of these resources are excellent.

Andrej Karpathy has very good video tutorials on how to write your own GPT: https://www.youtube.com/watch?v=kCc8FmEb1nY
This is excellent
3Blue1Brown has a pretty great video series walking through Transformers:

https://www.youtube.com/watch?v=wjZofJX0v4M

This paper is much clearer and succinct than the original papers from the field.

Trying to build a protein diffusion model from scratch right now.

The math explainer is quite helpful

The first diffusion model for text generation was published less than a year ago. Should add something about that
> see tutorial on diffusion > get excited > it's all math in latex > despair
Latex is great in this case, because the equations are all written out clearly.

If you want to understand diffusion, it's a little difficult to avoid math.

You may find the huggingface course more approachable

https://huggingface.co/learn/diffusion-course/en/unit0/1

Maybe check out the fast ai course. Jeremy Howard has a way of explaining stuff.
the math is learnable and with gpt nowadays, extremely so
You are right! LLMs accept the formulas and may explain them...

It will be more difficult to tell when they are wrong, though - when you cannot verify directly. But it will be a device for people to get acquainted with the math.

Diffusion is math. It’s difficult to avoid math if you want to understand math.
Be that as it may, this isn't what I'd call a "tutorial," in the sense that you'd better already have a strong command of the subject matter or you won't get much out of it.