about
Direction-Preserving Number Representations (arxiv.org)
1 point by matt_d 138 days ago | hide | past | pdf | discuss on HN

In plain words: They study how well a small set of values, used for each vector entry, can capture its direction, and find the best set. Standard formats like two's complement and floating point are provably worse, while NVIDIA's 4-bit format nearly matches the best possible set.

Abstract

Low-precision number formats are widely used in modern machine learning systems due to their efficiency. Accurate direction representation is key to the accuracy of vector operations. This work precisely explores the extent to which the direction of a vector can be represented by selecting its scalar elements from a common finite alphabet of a given size. This is standard practice in machine learning, where low-precision significands may be narrow-width floating-point or integer values. A geometric framework is introduced for analyzing the directional coverage of such product-structured codes. This work analytically quantifies the suboptimality gap between such product-structured codes and spherical codes for the vector as a whole, in both low and asymptotically high dimensions. Furthermore, within the product code class, it is proven that the standard formats of two's complement, fixed-point, and floating-point are suboptimal, again with quantified gap, pointing to the potential to develop new scalar number formats. Such scalar alphabets are numerically optimized across multiple block dimensions for directional coverage, including the dimension used in NVIDIA's NVFP4 format. Experimental results are presented comparing the performance of standard formats and the optimized alphabet. We find that for four bits, NVIDIA's choice of E2M1 closely approximates the optimized alphabet, providing a geometric explanation for its strong performance in low-precision machine learning workloads and an analytical understanding of the link between that superiority and block size. We provide open-source formal proofs in Lean for the theorems in this work, along with the experimental code and the optimized alphabets obtained.

Bardia Zadeh, George A. Constantinides
arXiv:2605.07662 · cs.LG, math.NA · submitted May 8, 2026
abstract · pdf · html · 9 pages excluding appendices and references, 18 in total. 5 figures

add comment on HN