about
Neural Networks with Few Multiplications (arxiv.org)
66 points by skybrian on Nov 9, 2015 | hide | past | pdf | 8 comments on HN

In plain words: To speed up training, weights are randomly set to plus or minus one and layer values rounded to powers of two, turning multiplications into sign flips and bit shifts. Across three image datasets, it matched or beat the accuracy of standard training.

Abstract

For most deep learning algorithms training is notoriously time consuming. Since most of the computation in training neural networks is typically spent on floating point multiplications, we investigate an approach to training that eliminates the need for most of these. Our method consists of two parts: First we stochastically binarize weights to convert multiplications involved in computing hidden states to sign changes. Second, while back-propagating error derivatives, in addition to binarizing the weights, we quantize the representations at each layer to convert the remaining multiplications into binary shifts. Experimental results across 3 popular datasets (MNIST, CIFAR10, SVHN) show that this approach not only does not hurt classification performance but can result in even better performance than standard stochastic gradient descent training, paving the way to fast, hardware-friendly training of neural networks.

Zhouhan Lin, Matthieu Courbariaux, Roland Memisevic, Yoshua Bengio
arXiv:1510.03009 · cs.LG, cs.NE · submitted Oct 11, 2015 · updated Feb 26, 2016
abstract · pdf · html · Published as a conference paper at ICLR 2016. 9 pages, 3 figures

add comment on HN
Also discussed: Oct 2015 (3 points, 0 comments)

I've seen lots of alternate approaches to constructing neural nets with varying operations and one of the common themes is that it doesn't really matter that much w.r.t. accuracy, so you may as well optimize for performance. This seems to indicate to me that neural nets are some variation on a yet unidentified more elegant mathematical structure that might enable better theoretical understanding about training phenomena. The spin-glass equivalencies I think further suggest more elegant underlying structures.

Once we understand that underlying structure we might be able to do really cool things, i.e. identify the nature and size of training data set required for solving a given problem, or train much, much faster.

Some decades ago, on one of the first IBM PCs 12Mhz, I created a neural network using static point multiplication (binary shifts) that learned to play tic tac toe over about a weekend, in C. I probably have some code for it still lying around.
I would be extremely interested in looking at that!
I would highly recommend this video https://www.youtube.com/watch?v=DleXA5ADG78 from Geoffrey Hinton in 2012 that I think is related to this research.
Wow. This opens a way for deep learning on much cheaper and less power-hungry hardware.
Does this not represent some "generalizable mechanism" for distributive elimination of floating point cycles or is this so specific to it's job that there is no generalization worth considering. My instinct says this may be somewhat generalizable making it significant in algorithms if it's actually working.
Yoshua Bengio is in this, you bet it's real!
Interesting, but the paper doesn't mention speed gains

It is definitely promising. Even though fp operations today are "cheap", integer ones are still faster