about
Flash Diffusion: Accelerating Any Conditional Diffusion Model (arxiv.org)
2 points by jasondavies on Jun 19, 2024 | hide | past | pdf | discuss on HN

In plain words: It teaches a pre-trained image generator to make pictures in a few steps instead of dozens, with a few hours of training and fewer adjustable parts than earlier shortcuts. It beat the best few-step results on standard image tests and handled text-to-image, editing, and upscaling.

Abstract · Flash Diffusion: Accelerating Any Conditional Diffusion Model for Few Steps Image Generation

In this paper, we propose an efficient, fast, and versatile distillation method to accelerate the generation of pre-trained diffusion models: Flash Diffusion. The method reaches state-of-the-art performances in terms of FID and CLIP-Score for few steps image generation on the COCO2014 and COCO2017 datasets, while requiring only several GPU hours of training and fewer trainable parameters than existing methods. In addition to its efficiency, the versatility of the method is also exposed across several tasks such as text-to-image, inpainting, face-swapping, super-resolution and using different backbones such as UNet-based denoisers (SD1.5, SDXL) or DiT (Pixart-$α$), as well as adapters. In all cases, the method allowed to reduce drastically the number of sampling steps while maintaining very high-quality image generation. The official implementation is available at https://github.com/gojasper/flash-diffusion.

Clément Chadebec, Onur Tasar, Eyal Benaroche, Benjamin Aubin
arXiv:2406.02347 · cs.CV, cs.AI, cs.LG · submitted Jun 4, 2024 · updated Dec 18, 2024
abstract · pdf · html · Accepted to AAAI 2025

add comment on HN