about
Matryoshka Quantization (arxiv.org)
2 points by fzliu on Feb 11, 2025 | hide | past | pdf | discuss on HN

In plain words: Weights are stored so the smallest 2-bit version sits inside the larger 8-bit one, like nesting dolls, letting one trained model serve at whatever precision a device needs. Its 2-bit versions beat the usual separately trained 2-bit models by up to 7%.

Abstract

Quantizing model weights is critical for reducing the communication and inference costs of large models. However, quantizing models -- especially to low precisions like int4 or int2 -- requires a trade-off in model quality; int2, in particular, is known to severely degrade model quality. Consequently, practitioners are often forced to maintain multiple models with different quantization levels or serve a single model that best satisfies the quality-latency trade-off. On the other hand, integer data types, such as int8, inherently possess a nested (Matryoshka) structure where smaller bit-width integers, like int4 or int2, are nested within the most significant bits. Leveraging this insight, in this paper, we propose Matryoshka Quantization (MatQuant), a novel multi-scale quantization technique that alleviates the aforementioned challenge. This technique allows us to train and maintain a single quantized model but serve it with the precision demanded by the deployment. Furthermore, leveraging MatQuant's co-training and co-distillation regularization, int2 precision models extracted by MatQuant outperform standard int2 quantization by up to to 4% and 7% with OmniQuant and QAT as base algorithms respectively. Finally, we demonstrate that by using an extra bit to represent outliers, a model with an effective precision of 2.05-bit gives an additional 6% improvement with OmniQuant as the base algorithm.

Pranav Nair, Puranjay Datta, Jeff Dean, Prateek Jain, Aditya Kusupati
arXiv:2502.06786 · cs.LG, cs.AI · submitted Feb 10, 2025 · updated Mar 3, 2025
abstract · pdf · html

add comment on HN
Also discussed: Feb 2025 (1 point, 0 comments) · Feb 2025 (2 points, 0 comments) · Feb 2025 (3 points, 0 comments) · Feb 2025 (4 points, 1 comment)