about
IFairy: The First 2-bit Complex LLM with All Parameters in \{\pm1, \pm i\} (arxiv.org)
4 points by Gathering6678 on Aug 16, 2025 | hide | past | pdf | 1 comment on HN

In plain words: Weights are stored as complex numbers with only four values, ±1 and ±i, so each takes 2 bits and needs only additions and swaps, not multiplication. It beats even the unquantized accuracy that today's 2-bit methods try to copy.

Abstract · iFairy: the First 2-bit Complex LLM with All Parameters in $\{\pm1, \pm i\}$

Quantization-Aware Training (QAT) integrates quantization into the training loop, enabling LLMs to learn robust low-bit representations, and is widely recognized as one of the most promising research directions. All current QAT research focuses on minimizing quantization error on full-precision models, where the full-precision accuracy acts as an upper bound (accuracy ceiling). No existing method has even attempted to surpass this ceiling. To break this ceiling, we propose a new paradigm: raising the ceiling (full-precision model), and then still quantizing it efficiently into 2 bits. We propose Fairy$\pm i$, the first 2-bit quantization framework for complex-valued LLMs. Specifically, our method leverages the representational advantages of the complex domain to boost full-precision accuracy. We map weights to the fourth roots of unity $\{\pm1, \pm i\}$, forming a perfectly symmetric and information-theoretically optimal 2-bit representation. Importantly, each quantized weight has either a zero real or imaginary part, enabling multiplication-free inference using only additions and element swaps. Experimental results show that Fairy$\pm i$ outperforms the ceiling of existing 2-bit quantization approaches in terms of both PPL and downstream tasks, while maintaining strict storage and compute efficiency. This work opens a new direction for building highly accurate and practical LLMs under extremely low-bit constraints.

Feiyu Wang, Guoan Wang, Yihao Zhang, Shengfan Wang, Weitao Li, Bokai Huang, Shimao Chen, Zihan Jiang, Rui Xu, Tong Yang
arXiv:2508.05571 · cs.LG, cs.CL · submitted Aug 7, 2025 · updated Aug 16, 2025
abstract · pdf · html · 15 pages, 9 figures

add comment on HN

This is great. I want to see them implement some of this in hardware, using at least FPGAs to build some core architecture to see how much area savings they get.