about
Double Descent and Malign Overfitting in Diffusion Models (arxiv.org)
2 points by Betelbuddy 10 days ago | hide | past | pdf | discuss on HN

In plain words: They tested why bigger image-generating models start memorizing training pictures instead of improving, using face-image experiments plus a simple math model. Quality worsens once parameters reach the number of training examples, but a penalty or early stopping makes big models beat every unregularized one.

Abstract

Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularization, interpolating models generalize well, the test error following a double-descent curve. One might expect the same benign overfitting for diffusion models, whose training reduces to regression, i.e. to minimizing a quadratic score-matching loss. Yet the opposite is observed: overfitting here is catastrophic, driving the model into a memorization regime. We resolve this paradox by combining experiments on U-Nets trained on CelebA with a random-features model for which we derive closed-form learning curves. We show that with a fixed number $m$ of noise realizations per training sample, an interpolation peak does occur, but at $p\sim nm$ rather than at $p\sim n$ as in standard regression. The rise of the test loss, however, sets in much earlier, at $p\sim n$, independently of $m$. This overfitting is malign because, although the implicit regularization of training is fully at work, it drives the model toward the empirical score, which memorizes the training set, rather than toward the true score. A bias-variance decomposition pinpoints the mechanism: the bias of the score estimator starts to grow at $p\sim n$; past the peak the variance decays, as in regression, whereas the bias keeps growing and both saturate at a large value. Since diffusion models are trained with $m\gg1$, the peak is pushed to very large model sizes, and therefore sit on the rising branch that precedes it, where malign overfitting is already in play. Nevertheless, overparameterization remains beneficial when paired with regularization: in the random-features theory and in U-Net experiments, optimally regularized large models---via a ridge penalty or early stopping, respectively---outperform any unregularized models.

Raphaël Urfin, Tony Bonnaire, Giulio Biroli, Marc Mézard
arXiv:2609.26392 · cs.LG, cond-mat.dis-nn · submitted Sep 22, 2026 · updated Sep 29, 2026
abstract · pdf · html · 44 pages, 17 figures

add comment on HN