about
Google's extreme AI compression paper was on arXiv since April 2025 (arxiv.org)
2 points by fadijob 192 days ago | hide | past | pdf | 1 comment on HN

In plain words: Vectors are randomly rotated so their coordinates act nearly independently, then each is compressed with a simple fixed quantizer, working online without training data. It comes within about 2.7 times of the theoretical best distortion, and at 3.5 bits per channel kept quality unchanged.

Abstract · TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate

Vector quantization, a problem rooted in Shannon's source coding theory, aims to quantize high-dimensional Euclidean vectors while minimizing distortion in their geometric structure. We propose TurboQuant to address both mean-squared error (MSE) and inner product distortion, overcoming limitations of existing methods that fail to achieve optimal distortion rates. Our data-oblivious algorithms, suitable for online applications, achieve near-optimal distortion rates (within a small constant factor) across all bit-widths and dimensions. TurboQuant achieves this by randomly rotating input vectors, inducing a concentrated Beta distribution on coordinates, and leveraging the near-independence property of distinct coordinates in high dimensions to simply apply optimal scalar quantizers per each coordinate. Recognizing that MSE-optimal quantizers introduce bias in inner product estimation, we propose a two-stage approach: applying an MSE quantizer followed by a 1-bit Quantized JL (QJL) transform on the residual, resulting in an unbiased inner product quantizer. We also provide a formal proof of the information-theoretic lower bounds on best achievable distortion rate by any vector quantizer, demonstrating that TurboQuant closely matches these bounds, differing only by a small constant ($\approx 2.7$) factor. Experimental results validate our theoretical findings, showing that for KV cache quantization, we achieve absolute quality neutrality with 3.5 bits per channel and marginal quality degradation with 2.5 bits per channel. Furthermore, in nearest neighbor search tasks, our method outperforms existing product quantization techniques in recall while reducing indexing time to virtually zero.

Amir Zandieh, Majid Daliri, Majid Hadian, Vahab Mirrokni
arXiv:2504.19874 · cs.LG, cs.AI, cs.DB, cs.DS · submitted Apr 28, 2025
abstract · pdf · html · 25 pages

add comment on HN

The arXiv paper was submitted April 2025, the research itself isn't new, but the new is Google's blog post packaging it for a wider audience.

worth reading the original paper alongside the blog post. I think the ppaper has details the blog post glosses over, particularly around the calibration-free quantization approach and how they handle outlier channels.

Interestingly: the research sits on arXiv for a year, nobody talks about it