In plain words: To speed up training, weights are randomly set to plus or minus one and layer values rounded to powers of two, turning multiplications into sign flips and bit shifts. Across three image datasets, it matched or beat the accuracy of standard training.
Abstract
For most deep learning algorithms training is notoriously time consuming. Since most of the computation in training neural networks is typically spent on floating point multiplications, we investigate an approach to training that eliminates the need for most of these. Our method consists of two parts: First we stochastically binarize weights to convert multiplications involved in computing hidden states to sign changes. Second, while back-propagating error derivatives, in addition to binarizing the weights, we quantize the representations at each layer to convert the remaining multiplications into binary shifts. Experimental results across 3 popular datasets (MNIST, CIFAR10, SVHN) show that this approach not only does not hurt classification performance but can result in even better performance than standard stochastic gradient descent training, paving the way to fast, hardware-friendly training of neural networks.
Zhouhan Lin, Matthieu Courbariaux, Roland Memisevic, Yoshua Bengio
arXiv:1510.03009 · cs.LG, cs.NE · submitted Oct 11, 2015 · updated Feb 26, 2016
abstract · pdf · html · Published as a conference paper at ICLR 2016. 9 pages, 3 figures
Once we understand that underlying structure we might be able to do really cool things, i.e. identify the nature and size of training data set required for solving a given problem, or train much, much faster.