about
Unified Latents (UL): How to train your latents (arxiv.org)
2 points by pama 223 days ago | hide | past | pdf | discuss on HN

In plain words: A system learns a compact code by training the encoder, a noise prior, and decoder together, matching its noise to the prior's lowest level to keep it small. It scored 1.4 on ImageNet-512 with less training compute than Stable Diffusion's code, and 1.3 on video.

Abstract

We present Unified Latents (UL), a framework for learning latent representations that are jointly regularized by a diffusion prior and decoded by a diffusion model. By linking the encoder's output noise to the prior's minimum noise level, we obtain a simple training objective that provides a tight upper bound on the latent bitrate. On ImageNet-512, our approach achieves competitive FID of 1.4, with high reconstruction quality (PSNR) while requiring fewer training FLOPs than models trained on Stable Diffusion latents. On Kinetics-600, we set a new state-of-the-art FVD of 1.3.

Jonathan Heek, Emiel Hoogeboom, Thomas Mensink, Tim Salimans
arXiv:2602.17270 · cs.LG, cs.CV · submitted Feb 19, 2026
abstract · pdf · html

add comment on HN