In plain words: They compared three ways of zeroing out unimportant weights in large translation and image-recognition models, across thousands of runs. Dropping the smallest weights matched or beat fancier approaches, and the networks it left could not be retrained from scratch to the same accuracy.
Abstract · The State of Sparsity in Deep Neural Networks
We rigorously evaluate three state-of-the-art techniques for inducing sparsity in deep neural networks on two large-scale learning tasks: Transformer trained on WMT 2014 English-to-German, and ResNet-50 trained on ImageNet. Across thousands of experiments, we demonstrate that complex techniques (Molchanov et al., 2017; Louizos et al., 2017b) shown to yield high compression rates on smaller datasets perform inconsistently, and that simple magnitude pruning approaches achieve comparable or better results. Additionally, we replicate the experiments performed by (Frankle & Carbin, 2018) and (Liu et al., 2018) at scale and show that unstructured sparse architectures learned through pruning cannot be trained from scratch to the same test set performance as a model trained with joint sparsification and optimization. Together, these results highlight the need for large-scale benchmarks in the field of model compression. We open-source our code, top performing model checkpoints, and results of all hyperparameter configurations to establish rigorous baselines for future work on compression and sparsification.
Trevor Gale, Erich Elsen, Sara Hooker
arXiv:1902.09574 · cs.LG, stat.ML · submitted Feb 25, 2019
abstract · pdf · html
My biggest question coming out of this work was as follows: which small scale (or - at the very least - inexpensive) benchmarks share enough properties in common with these large scale networks that we should expect results to scale with reasonable fidelity? Resnet50 is still far too slow and expensive to use as a day-to-day research network in academia, let alone transformer. Personally, I've found resnet18 on CIFAR10 to pretty reliably predict behavior on resnet50 on imagenet, but that's anecdotal. For the academics who can't drop hundreds of thousands of dollars (or more) on each paper but still want to contribute to research progress, we should carefully assess (or design) benchmarks with this property in mind.
(With respect to the lottery ticket hypothesis, we have a complimentary ICML submission about its behavior on large-scale networks coming shortly!)