about
Scalable High-Resolution Pixel-Space Image Synthesis (arxiv.org)
1 point by artninja1988 on Jan 23, 2024 | hide | past | pdf | discuss on HN

In plain words: A transformer that shrinks the image to work at coarse sizes then expands it again keeps cost growing in step with pixel count, letting it generate 1024×1024 images from pixels. It matched leading ImageNet models and set the best diffusion result on FFHQ-1024 faces.

Abstract · Scalable High-Resolution Pixel-Space Image Synthesis with Hourglass Diffusion Transformers

We present the Hourglass Diffusion Transformer (HDiT), an image generative model that exhibits linear scaling with pixel count, supporting training at high-resolution (e.g. $1024 \times 1024$) directly in pixel-space. Building on the Transformer architecture, which is known to scale to billions of parameters, it bridges the gap between the efficiency of convolutional U-Nets and the scalability of Transformers. HDiT trains successfully without typical high-resolution training techniques such as multiscale architectures, latent autoencoders or self-conditioning. We demonstrate that HDiT performs competitively with existing models on ImageNet $256^2$, and sets a new state-of-the-art for diffusion models on FFHQ-$1024^2$.

Katherine Crowson, Stefan Andreas Baumann, Alex Birch, Tanishq Mathew Abraham, Daniel Z. Kaplan, Enrico Shippole
arXiv:2401.11605 · cs.CV, cs.AI, cs.LG · submitted Jan 21, 2024 · updated Mar 25, 2026
abstract · pdf · html · 20 pages, 13 figures, project page and code available at https://crowsonkb.github.io/hourglass-diffusion-transformers/

add comment on HN