about
Emergence of Diffusion Models from Associative Memory (arxiv.org)
9 points by fzliu on Jun 20, 2025 | hide | past | pdf | 2 comments on HN

In plain words: Viewing image generators as memory-retrieval systems, the study finds they store each training example as its own attractor when data is scarce. As data grows, new attractors appear that match no training example — the first sign of real generation, across many architectures and datasets.

Abstract · Memorization to Generalization: Emergence of Diffusion Models from Associative Memory

Dense Associative Memories (DenseAMs) are generalizations of Hopfield networks, which have superior information storage capacity and can store training data points (memories) at local minima of the energy landscape. When the amount of training data exceeds the critical memory storage capacity of these models, new local minima, which are different from the training data, emerge. In Associative Memory these emergent local minima are called $\textit{spurious}\; \textit{states}$, which hinder memory retrieval. In this work, we examine diffusion models (DMs) through the DenseAM lens, viewing their generative process as an attempt of a memory retrieval. In the small data regimes, DMs create distinct attractors for each training sample, akin to DenseAMs below the critical memory storage. As the training data size increases, they transition from memorization to generalization. We identify a critical intermediate phase, predicted by DenseAM theory -- the spurious states. In generative modeling, these states are no longer negative artifacts but rather are the first signs of generative capabilities. We characterize the basins of attraction, energy landscape curvature, and computational properties of these previously overlooked states. Their existence is demonstrated across a wide range of architectures and datasets.

Bao Pham, Gabriel Raya, Matteo Negri, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov
arXiv:2505.21777 · cs.LG, cond-mat.dis-nn, cs.CV, q-bio.NC, stat.ML · submitted May 27, 2025 · updated Mar 16, 2026
abstract · pdf · html

add comment on HN

The paper maps two energy-based methods to each other. I'm not sure what the innovation is.
The novelty is in identifying spurious samples at the memorization-generalization threshold that act like memorized samples even though they don't appear in the training data.