about
Fit: Flexible Vision Transformer for Diffusion Model (arxiv.org)
3 points by dihuang on Feb 20, 2024 | hide | past | pdf | 2 comments on HN

In plain words: This image generator treats a picture as a flexible string of tokens that can grow or shrink, instead of a fixed grid, so it can make images at any size or shape. It stayed strong across many resolutions, including sizes it never trained on, where fixed-grid models struggle.

Abstract · FiT: Flexible Vision Transformer for Diffusion Model

Nature is infinitely resolution-free. In the context of this reality, existing diffusion models, such as Diffusion Transformers, often face challenges when processing image resolutions outside of their trained domain. To overcome this limitation, we present the Flexible Vision Transformer (FiT), a transformer architecture specifically designed for generating images with unrestricted resolutions and aspect ratios. Unlike traditional methods that perceive images as static-resolution grids, FiT conceptualizes images as sequences of dynamically-sized tokens. This perspective enables a flexible training strategy that effortlessly adapts to diverse aspect ratios during both training and inference phases, thus promoting resolution generalization and eliminating biases induced by image cropping. Enhanced by a meticulously adjusted network structure and the integration of training-free extrapolation techniques, FiT exhibits remarkable flexibility in resolution extrapolation generation. Comprehensive experiments demonstrate the exceptional performance of FiT across a broad range of resolutions, showcasing its effectiveness both within and beyond its training resolution distribution. Repository available at https://github.com/whlzy/FiT.

Zeyu Lu, Zidong Wang, Di Huang, Chengyue Wu, Xihui Liu, Wanli Ouyang, Lei Bai
arXiv:2402.12376 · cs.CV · submitted Feb 19, 2024 · updated Oct 15, 2024
abstract · pdf · html

add comment on HN
Also discussed: Feb 2024 (1 point, 0 comments)

Hadn't heard of folks using 2D RoPE before. Is that new work here, or an established technique?

(The way the authors present it, makes it seem like they are introducing it as a concept, but they never outright say they came up with it, AFAICT)

I didn't go through all literature, but it's possible that some other team focused on 1D, made one toy experiment with 2D just for giggles. That happens a lot, but it's just my speculation that this happened in this case.