about
Movement Pruning: Adaptive Sparsity by Fine-Tuning (arxiv.org)
1 point by blopeur on May 19, 2020 | hide | past | pdf | discuss on HN

In plain words: When fine-tuning a pretrained language model, it deletes weights shrinking toward zero instead of the smallest ones, so cuts follow training. Combined with a teacher model's guidance, it kept 3% of the parameters with little accuracy loss, beating size-based pruning especially at high sparsity.

Abstract

Magnitude pruning is a widely used strategy for reducing model size in pure supervised learning; however, it is less effective in the transfer learning regime that has become standard for state-of-the-art natural language processing applications. We propose the use of movement pruning, a simple, deterministic first-order weight pruning method that is more adaptive to pretrained model fine-tuning. We give mathematical foundations to the method and compare it to existing zeroth- and first-order pruning methods. Experiments show that when pruning large pretrained language models, movement pruning shows significant improvements in high-sparsity regimes. When combined with distillation, the approach achieves minimal accuracy loss with down to only 3% of the model parameters.

Victor Sanh, Thomas Wolf, Alexander M. Rush
arXiv:2005.07683 · cs.CL, cs.LG · submitted May 15, 2020 · updated Oct 23, 2020
abstract · pdf · html · 14 pages, 6 figures, 3 tables. Published at NeurIPS2020. Code: \url{huggingface.co/mvp}

add comment on HN
Also discussed: Jun 2020 (1 point, 0 comments)