In plain words: An image generator trained without depth labels still builds internal signals for how far away things are and which parts are the main object. Simple tests read these signals early in noisy cleanup steps, and nudging them changes the picture, showing the model uses them.
Abstract · Beyond Surface Statistics: Scene Representations in a Latent Diffusion Model
Latent diffusion models (LDMs) exhibit an impressive ability to produce realistic images, yet the inner workings of these models remain mysterious. Even when trained purely on images without explicit depth information, they typically output coherent pictures of 3D scenes. In this work, we investigate a basic interpretability question: does an LDM create and use an internal representation of simple scene geometry? Using linear probes, we find evidence that the internal activations of the LDM encode linear representations of both 3D depth data and a salient-object / background distinction. These representations appear surprisingly early in the denoising process$-$well before a human can easily make sense of the noisy images. Intervention experiments further indicate these representations play a causal role in image synthesis, and may be used for simple high-level editing of an LDM's output. Project page: https://yc015.github.io/scene-representation-diffusion-model/
Yida Chen, Fernanda Viégas, Martin Wattenberg
arXiv:2306.05720 · cs.CV, cs.AI, cs.LG · submitted Jun 9, 2023 · updated Nov 4, 2023
abstract · pdf · html · A short version of this paper is accepted in the NeurIPS 2023 Workshop on Diffusion Models: https://nips.cc/virtual/2023/74894
> Researchers experimentally discovered that image-generating AI Stable Diffusion v1 uses internal representations of 3D geometry - depth maps and object saliency maps - when generating an image. This ability emerged during the training phase of the AI, and was not programmed by people.
https://www.reddit.com/r/StableDiffusion/comments/15wvz2a/re...