about
Low-Memory Neural Network Training: A Technical Report (arxiv.org)
58 points by wildermuthn on Sep 21, 2019 | hide | past | pdf | 3 comments on HN

In plain words: They measured how much memory training needs on image and translation models, then tested four ways to shrink it: sparse weights, fewer-decimal numbers, smaller batches, and recomputing work instead of storing it. Together they cut image-model training memory 60-fold with little accuracy loss.

Abstract

Memory is increasingly often the bottleneck when training neural network models. Despite this, techniques to lower the overall memory requirements of training have been less widely studied compared to the extensive literature on reducing the memory requirements of inference. In this paper we study a fundamental question: How much memory is actually needed to train a neural network? To answer this question, we profile the overall memory usage of training on two representative deep learning benchmarks -- the WideResNet model for image classification and the DynamicConv Transformer model for machine translation -- and comprehensively evaluate four standard techniques for reducing the training memory requirements: (1) imposing sparsity on the model, (2) using low precision, (3) microbatching, and (4) gradient checkpointing. We explore how each of these techniques in isolation affects both the peak memory usage of training and the quality of the end model, and explore the memory, accuracy, and computation tradeoffs incurred when combining these techniques. Using appropriate combinations of these techniques, we show that it is possible to the reduce the memory required to train a WideResNet-28-2 on CIFAR-10 by up to 60.7x with a 0.4% loss in accuracy, and reduce the memory required to train a DynamicConv model on IWSLT'14 German to English translation by up to 8.7x with a BLEU score drop of 0.15.

Nimit S. Sohoni, Christopher R. Aberger, Megan Leszczynski, Jian Zhang, Christopher Ré
arXiv:1904.10631 · cs.LG, stat.ML · submitted Apr 24, 2019 · updated Apr 8, 2022
abstract · pdf · html · Version notes: Copyedits and citation fixes

add comment on HN
Also discussed: Apr 2019 (8 points, 0 comments)

This it the best explanation I've read of how memory is allocated to training deep learning models, and what possible solutions there are to reducing that footprint.

Some things I already knew about, such as gradient checkpointing and FP16. What was new to me is microbatching (gradient accumulation).

Many of the large models appearing these days (Transformers in particular) are really costly to train from scratch. What I've noticed, about BERT in particular, is that none of the memory-saving techniques are used. I suppose a large corporation doesn't mind spending more money on compute, but for a startup, these techniques could be quite useful on a limited budget. The cost to accuracy appears small, given the right mix of methods.

note: Xlnet (successor of BERT) has a pull request allowing FP16 use.

And yes it's sad that Google don't prioritize this. More generally, researchers works in isolation, they don't integrate other researchers synergetic ideas...

I feel like Tensorflow tends to allocate scratch space in GPU RAM for the result of every math operation.

So if you have a big tensor and want to do math on it, you need space to effectively store multiple copies of the tensor.

I think TF does this for speed reasons.