about
Elastic Looped Transformers for Visual Generation (arxiv.org)
1 point by gmays 171 days ago | hide | past | pdf | discuss on HN

In plain words: Instead of stacking layers, this visual generator runs one transformer block repeatedly, training shorter loop counts to match the longest so any depth works. It uses four times fewer parameters than deep stacks at the same computing cost, with an ImageNet quality score of 2.0.

Abstract · ELT: Elastic Looped Transformers for Visual Generation

We introduce Elastic Looped Transformers (ELT), a highly parameter-efficient class of visual generative models based on a recurrent transformer architecture. While conventional generative models rely on deep stacks of unique transformer layers, our approach employs iterative, weight-shared transformer blocks to drastically reduce parameter counts while maintaining high synthesis quality. To effectively train these models for image and video generation, we propose the idea of Intra-Loop Self Distillation (ILSD), where student configurations (intermediate loops) are distilled from the teacher configuration (maximum training loops) to ensure consistency across the model's depth in a single training step. Our framework yields a family of elastic models from a single training run, enabling Any-Time inference capability with dynamic trade-offs between computational cost and generation quality, with the same parameter count. ELT significantly shifts the efficiency frontier for visual synthesis. With $4\times$ reduction in parameter count under iso-inference-compute settings, ELT achieves a competitive FID of $2.0$ on class-conditional ImageNet $256 \times 256$ and FVD of $72.8$ on class-conditional UCF-101.

Sahil Goyal, Swayam Agrawal, Gautham Govind Anil, Prateek Jain, Sujoy Paul, Aditya Kusupati
arXiv:2604.09168 · cs.CV · submitted Apr 10, 2026 · updated Jul 23, 2026
abstract · pdf · html · Accepted to ECCV 2026

add comment on HN