about
Towards Trillion Parameter Language Model with Sparse Heterogeneous Computing (arxiv.org)
1 point by mlejva on Mar 21, 2023 | hide | past | pdf | discuss on HN

In plain words: A trillion-parameter language model that randomly activates only a few expert parts per word, keeping those experts on separate chips from the ones doing the math. This split sped up training 6.3 times and gave the best zero-shot results on Chinese language tasks.

Abstract · PanGu-Σ: Towards Trillion Parameter Language Model with Sparse Heterogeneous Computing

The scaling of large language models has greatly improved natural language understanding, generation, and reasoning. In this work, we develop a system that trained a trillion-parameter language model on a cluster of Ascend 910 AI processors and MindSpore framework, and present the language model with 1.085T parameters named PanGu-Σ. With parameter inherent from PanGu-α, we extend the dense Transformer model to sparse one with Random Routed Experts (RRE), and efficiently train the model over 329B tokens by using Expert Computation and Storage Separation(ECSS). This resulted in a 6.3x increase in training throughput through heterogeneous computing. Our experimental findings show that PanGu-Σ provides state-of-the-art performance in zero-shot learning of various Chinese NLP downstream tasks. Moreover, it demonstrates strong abilities when fine-tuned in application data of open-domain dialogue, question answering, machine translation and code generation.

Xiaozhe Ren, Pingyi Zhou, Xinfan Meng, Xinjing Huang, Yadao Wang, Weichao Wang, Pengfei Li, Xiaoda Zhang, Alexander Podolskiy, Grigory Arshinov, Andrey Bout, Irina Piontkovskaya, et al.
arXiv:2303.10845 · cs.CL · submitted Mar 20, 2023
abstract · pdf · html

add comment on HN