about
Instant Quantization of NN's Using Monte Carlo (arxiv.org)
2 points by Katydid on Jun 4, 2019 | hide | past | pdf | discuss on HN

In plain words: Instead of retraining a network to fit small integer numbers, it randomly samples each weight and activation to pick a low-bit value, making the network smaller and cheaper to run. It loses almost no accuracy and matches or beats methods that need extra training.

Abstract · Instant Quantization of Neural Networks using Monte Carlo Methods

Low bit-width integer weights and activations are very important for efficient inference, especially with respect to lower power consumption. We propose Monte Carlo methods to quantize the weights and activations of pre-trained neural networks without any re-training. By performing importance sampling we obtain quantized low bit-width integer values from full-precision weights and activations. The precision, sparsity, and complexity are easily configurable by the amount of sampling performed. Our approach, called Monte Carlo Quantization (MCQ), is linear in both time and space, with the resulting quantized, sparse networks showing minimal accuracy loss when compared to the original full-precision networks. Our method either outperforms or achieves competitive results on multiple benchmarks compared to previous quantization methods that do require additional training.

Gonçalo Mordido, Matthijs Van Keirsbilck, Alexander Keller
arXiv:1905.12253 · cs.LG, stat.ML · submitted May 29, 2019 · updated Jan 7, 2020
abstract · pdf · html

add comment on HN