about
8-Bit Numerical Formats for Deep Neural Networks (2022) (arxiv.org)
1 point by ashvardanian on Jul 18, 2023 | hide | past | pdf | 1 comment on HN

In plain words: Storing weights, activations and gradients as 8-bit floating-point numbers instead of 32-bit, the study tests different splits of exponent and precision bits to find the best layout. The right layout trained image and language models faster and used less power with no accuracy loss.

Abstract · 8-bit Numerical Formats for Deep Neural Networks

Given the current trend of increasing size and complexity of machine learning architectures, it has become of critical importance to identify new approaches to improve the computational efficiency of model training. In this context, we address the advantages of floating-point over fixed-point representation, and present an in-depth study on the use of 8-bit floating-point number formats for activations, weights, and gradients for both training and inference. We explore the effect of different bit-widths for exponents and significands and different exponent biases. The experimental results demonstrate that a suitable choice of these low-precision formats enables faster training and reduced power consumption without any degradation in accuracy for a range of deep learning models for image classification and language processing.

Badreddine Noune, Philip Jones, Daniel Justus, Dominic Masters, Carlo Luschi
arXiv:2206.02915 · cs.LG · submitted Jun 6, 2022
abstract · pdf · html

add comment on HN
Also discussed: Jun 2022 (2 points, 0 comments)

I am curious if there is anyone familiar with the standardization progress of those 8-bit representations?

PS: Assuming they can only have 256 unique values, would have been great if the papers provided conversion tables for every 8-bit float to a 32-bit value, to make testing accuracy degradation easier across applications.