about
Transfusion: Predict the next token and diffuse images with one multimodal model (arxiv.org)
122 points by fzliu on Sep 9, 2024 | hide | past | pdf | 10 comments on HN

In plain words: One transformer learns text and images together: it predicts the next word for text and cleans up noisy pixels step by step for pictures. It scaled better than chopping images into word-like tokens, and at 7 billion parameters matched dedicated text and image generators.

Abstract · Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

We introduce Transfusion, a recipe for training a multi-modal model over discrete and continuous data. Transfusion combines the language modeling loss function (next token prediction) with diffusion to train a single transformer over mixed-modality sequences. We pretrain multiple Transfusion models up to 7B parameters from scratch on a mixture of text and image data, establishing scaling laws with respect to a variety of uni- and cross-modal benchmarks. Our experiments show that Transfusion scales significantly better than quantizing images and training a language model over discrete image tokens. By introducing modality-specific encoding and decoding layers, we can further improve the performance of Transfusion models, and even compress each image to just 16 patches. We further demonstrate that scaling our Transfusion recipe to 7B parameters and 2T multi-modal tokens produces a model that can generate images and text on a par with similar scale diffusion models and language models, reaping the benefits of both worlds.

Chunting Zhou, Lili Yu, Arun Babu, Kushal Tirumala, Michihiro Yasunaga, Leonid Shamis, Jacob Kahn, Xuezhe Ma, Luke Zettlemoyer, Omer Levy
arXiv:2408.11039 · cs.AI, cs.CV · submitted Aug 20, 2024
abstract · pdf · html · 23 pages

add comment on HN
Also discussed: Aug 2024 (1 point, 0 comments)

This is such a natural extension to LLMs. I’m shocked it hasn’t been tried before.

When I ask a diffusion model to generate a chessboard, I’d expect the pieces to be placed randomly. We are getting closer to image generators that not only know what chess pieces look like but also where to place them.

You can talk to the authors directly on alphaXiv! https://www.alphaxiv.org/abs/2408.11039v1
It doesn't look like they are active there.
Stupid question: is their 7B model available? Is there public inference code that we could run? Or do they not usually release them along with these kinds of papers?
Doesn't appear to be any weights uploaded anywhere that I can find.

There are the starts of two (non-original-author) public implementations available on Github, but again -- doesn't appear to be any pretrained weights in either.

* https://github.com/lucidrains/transfusion-pytorch

* https://github.com/VachanVY/Transfusion.torch

I’d also like to know this.
Hmm. I wonder if this is similar to Diffusion Transformers?
this is somewhat similar, but diffusion transformers typically use a pre-trained text model as the text conditioning whereas, in this case it's integrated and trained together multimodally.
Would such a model be able to give more accurate description of images as well?
I think so, specially with finetuning