about
Mixture of a Million Experts (arxiv.org)
6 points by fofoz on Jul 9, 2024 | hide | past | pdf | discuss on HN

In plain words: A new transformer layer picks a few tiny experts from a pool of over a million, using a lookup so only the chosen ones run. It beat plain dense layers and usual mixtures with fewer, bigger experts on performance for the same computing cost.

Abstract · Mixture of A Million Experts

The feedforward (FFW) layers in standard transformer architectures incur a linear increase in computational costs and activation memory as the hidden layer width grows. Sparse mixture-of-experts (MoE) architectures have emerged as a viable approach to address this issue by decoupling model size from computational cost. The recent discovery of the fine-grained MoE scaling law shows that higher granularity leads to better performance. However, existing MoE models are limited to a small number of experts due to computational and optimization challenges. This paper introduces PEER (parameter efficient expert retrieval), a novel layer design that utilizes the product key technique for sparse retrieval from a vast pool of tiny experts (over a million). Experiments on language modeling tasks demonstrate that PEER layers outperform dense FFWs and coarse-grained MoEs in terms of performance-compute trade-off. By enabling efficient utilization of a massive number of experts, PEER unlocks the potential for further scaling of transformer models while maintaining computational efficiency.

Xu Owen He
arXiv:2407.04153 · cs.LG, cs.AI · submitted Jul 4, 2024
abstract · pdf · html

add comment on HN
Also discussed: Jul 2024 (2 points, 0 comments) · Jul 2024 (4 points, 0 comments) · Jul 2024 (3 points, 0 comments)