about
Extracting Training Data from Diffusion Models (arxiv.org)
3 points by tbruckner on Jan 31, 2023 | hide | past | pdf | discuss on HN

In plain words: By generating many images and keeping only ones that nearly match real training pictures, the team pulled over a thousand training images out of leading image generators. These models leak far more training data than older generators that draw pictures in one step.

Abstract

Image diffusion models such as DALL-E 2, Imagen, and Stable Diffusion have attracted significant attention due to their ability to generate high-quality synthetic images. In this work, we show that diffusion models memorize individual images from their training data and emit them at generation time. With a generate-and-filter pipeline, we extract over a thousand training examples from state-of-the-art models, ranging from photographs of individual people to trademarked company logos. We also train hundreds of diffusion models in various settings to analyze how different modeling and data decisions affect privacy. Overall, our results show that diffusion models are much less private than prior generative models such as GANs, and that mitigating these vulnerabilities may require new advances in privacy-preserving training.

Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian Tramèr, Borja Balle, Daphne Ippolito, Eric Wallace
arXiv:2301.13188 · cs.CR, cs.CV, cs.LG · submitted Jan 30, 2023
abstract · pdf · html

add comment on HN
Also discussed: Jan 2023 (163 points, 309 comments)