In plain words: Instead of rebuilding an image in one step from its compressed code, this system starts with noise and repeatedly cleans it up, using the code as a guide. That gave 22% better generation at the same compression, or 2.3x faster results with tighter compression.
Abstract · Epsilon-VAE: Denoising as Visual Decoding
In generative modeling, tokenization simplifies complex data into compact, structured representations, creating a more efficient, learnable space. For high-dimensional visual data, it reduces redundancy and emphasizes key features for high-quality generation. Current visual tokenization methods rely on a traditional autoencoder framework, where the encoder compresses data into latent representations, and the decoder reconstructs the original input. In this work, we offer a new perspective by proposing denoising as decoding, shifting from single-step reconstruction to iterative refinement. Specifically, we replace the decoder with a diffusion process that iteratively refines noise to recover the original image, guided by the latents provided by the encoder. We evaluate our approach by assessing both reconstruction (rFID) and generation quality (FID), comparing it to state-of-the-art autoencoding approaches. By adopting iterative reconstruction through diffusion, our autoencoder, namely Epsilon-VAE, achieves high reconstruction quality, which in turn enhances downstream generation quality by 22% at the same compression rates or provides 2.3x inference speedup through increasing compression rates. We hope this work offers new insights into integrating iterative generation and autoencoding for improved compression and generation.
Long Zhao, Sanghyun Woo, Ziyu Wan, Yandong Li, Han Zhang, Boqing Gong, Hartwig Adam, Xuhui Jia, Ting Liu
arXiv:2410.04081 · cs.CV, cs.AI, eess.IV · submitted Oct 5, 2024 · updated May 28, 2025
abstract · pdf · html · Accepted to ICML 2025. v2: added comparisons to SD-VAE and more visual results; v3: minor change to title; v4: camera-ready version
I am surprised that this works at high compression rates.
I would have thought that the more you squeeze into a latent the less correlated the individual latent values become. If they are correlated you can store them more efficiently by having the decoder know about the correlation and store the variance from that correlation with more precision. If that happens, and the values are uncorrelated then bilinear sampling would surely be counterproductive.
I feel like even a tiny-brained traditional VAE decoder would be able to do a better job at transforming the latent into a good conditioning for the U-Net.