about
aMUSEd: A lightweight masked-image-model (MIM) intended for fast generation (arxiv.org)
2 points by suraj_p on Jan 4, 2024 | hide | past | pdf | 1 comment on HN

In plain words: Instead of cleaning up noise step by step like today's text-to-image tools, this model fills in masked-out patches of an image from the text prompt. It uses a tenth of the original's size, needs fewer steps, and can learn a new style from one picture.

Abstract · aMUSEd: An Open MUSE Reproduction

We present aMUSEd, an open-source, lightweight masked image model (MIM) for text-to-image generation based on MUSE. With 10 percent of MUSE's parameters, aMUSEd is focused on fast image generation. We believe MIM is under-explored compared to latent diffusion, the prevailing approach for text-to-image generation. Compared to latent diffusion, MIM requires fewer inference steps and is more interpretable. Additionally, MIM can be fine-tuned to learn additional styles with only a single image. We hope to encourage further exploration of MIM by demonstrating its effectiveness on large-scale text-to-image generation and releasing reproducible training code. We also release checkpoints for two models which directly produce images at 256x256 and 512x512 resolutions.

Suraj Patil, William Berman, Robin Rombach, Patrick von Platen
arXiv:2401.01808 · cs.CV · submitted Jan 3, 2024
abstract · pdf · html

add comment on HN