about
DeepNet: Scaling Transformers to 1k Layers (arxiv.org)
1 point by ashvardanian on Mar 2, 2022 | hide | past | pdf | discuss on HN

In plain words: Rescaling each skip connection and choosing starting weights carefully keeps the network's updates from blowing up as layers pile up. It trained 1,000 layers easily, and a 200-layer version beat a 48-layer model nearly four times its size by 5 translation-score points.

Abstract · DeepNet: Scaling Transformers to 1,000 Layers

In this paper, we propose a simple yet effective method to stabilize extremely deep Transformers. Specifically, we introduce a new normalization function (DeepNorm) to modify the residual connection in Transformer, accompanying with theoretically derived initialization. In-depth theoretical analysis shows that model updates can be bounded in a stable way. The proposed method combines the best of two worlds, i.e., good performance of Post-LN and stable training of Pre-LN, making DeepNorm a preferred alternative. We successfully scale Transformers up to 1,000 layers (i.e., 2,500 attention and feed-forward network sublayers) without difficulty, which is one order of magnitude deeper than previous deep Transformers. Remarkably, on a multilingual benchmark with 7,482 translation directions, our 200-layer model with 3.2B parameters significantly outperforms the 48-layer state-of-the-art model with 12B parameters by 5 BLEU points, which indicates a promising scaling direction.

Hongyu Wang, Shuming Ma, Li Dong, Shaohan Huang, Dongdong Zhang, Furu Wei
arXiv:2203.00555 · cs.CL, cs.LG · submitted Mar 1, 2022
abstract · pdf · html · Work in progress

add comment on HN
Also discussed: Mar 2022 (194 points, 38 comments)