about
1.58-Bit Flux (arxiv.org)
2 points by reynaldi on Jan 1, 2025 | hide | past | pdf | 1 comment on HN

In plain words: A text-to-image generator's weights are squeezed into three values—minus one, zero, or plus one—using only the model's own signals, without any image data. It keeps 1024-pixel quality nearly the same while cutting storage 7.7 times and memory 5.1 times versus the original.

Abstract · 1.58-bit FLUX

We present 1.58-bit FLUX, the first successful approach to quantizing the state-of-the-art text-to-image generation model, FLUX.1-dev, using 1.58-bit weights (i.e., values in {-1, 0, +1}) while maintaining comparable performance for generating 1024 x 1024 images. Notably, our quantization method operates without access to image data, relying solely on self-supervision from the FLUX.1-dev model. Additionally, we develop a custom kernel optimized for 1.58-bit operations, achieving a 7.7x reduction in model storage, a 5.1x reduction in inference memory, and improved inference latency. Extensive evaluations on the GenEval and T2I Compbench benchmarks demonstrate the effectiveness of 1.58-bit FLUX in maintaining generation quality while significantly enhancing computational efficiency.

Chenglin Yang, Celong Liu, Xueqing Deng, Dongwon Kim, Xing Mei, Xiaohui Shen, Liang-Chieh Chen
arXiv:2412.18653 · cs.CV, cs.AI, cs.LG · submitted Dec 24, 2024
abstract · pdf · html

add comment on HN

Some impressive results

> 1.58-bit FLUX achieves a 7.7× reduction in model storage and more than a 5.1× reduction in inference memory usage