about
A Survey on Transformer Compression (arxiv.org)
1 point by PaulHoule on Feb 19, 2024 | hide | past | pdf | discuss on HN

In plain words: It collects ways to shrink Transformer models for everyday devices, grouped into four kinds: dropping weak parts, using fewer bits, teaching a small model from a big one, and redesigning layers. It compares them for language and vision and lays out open questions.

Abstract

Transformer plays a vital role in the realms of natural language processing (NLP) and computer vision (CV), specially for constructing large language models (LLM) and large vision models (LVM). Model compression methods reduce the memory and computational cost of Transformer, which is a necessary step to implement large language/vision models on practical devices. Given the unique architecture of Transformer, featuring alternative attention and feedforward neural network (FFN) modules, specific compression techniques are usually required. The efficiency of these compression methods is also paramount, as retraining large models on the entire training dataset is usually impractical. This survey provides a comprehensive review of recent compression methods, with a specific focus on their application to Transformer-based models. The compression methods are primarily categorized into pruning, quantization, knowledge distillation, and efficient architecture design (Mamba, RetNet, RWKV, etc.). In each category, we discuss compression methods for both language and vision tasks, highlighting common underlying principles. Finally, we delve into the relation between various compression methods, and discuss further directions in this domain.

Yehui Tang, Yunhe Wang, Jianyuan Guo, Zhijun Tu, Kai Han, Hailin Hu, Dacheng Tao
arXiv:2402.05964 · cs.LG, cs.CL, cs.CV · submitted Feb 5, 2024 · updated Apr 7, 2024
abstract · pdf · html · Model Compression, Transformer, Large Language Model, Large Vision Model, LLM

add comment on HN