In plain words: They mix Mamba, a fast text model that carries a running summary instead of re-reading everything, with specialized sub-networks where only a few handle each word. It matched plain Mamba's quality in 2.35 times fewer training steps and beat the usual Transformer version.
Abstract
State Space Models (SSMs) have become serious contenders in the field of sequential modeling, challenging the dominance of Transformers. At the same time, Mixture of Experts (MoE) has significantly improved Transformer-based Large Language Models, including recent state-of-the-art open models. We propose that to unlock the potential of SSMs for scaling, they should be combined with MoE. We showcase this on Mamba, a recent SSM-based model that achieves remarkable performance. Our model, MoE-Mamba, outperforms both Mamba and baseline Transformer-MoE. In particular, MoE-Mamba reaches the same performance as Mamba in $2.35\times$ fewer training steps while preserving the inference performance gains of Mamba against Transformer.
Maciej Pióro, Kamil Ciebiera, Krystian Król, Jan Ludziejewski, Michał Krutul, Jakub Krajewski, Szymon Antoniak, Piotr Miłoś, Marek Cygan, Sebastian Jaszczur
arXiv:2401.04081 · cs.LG, cs.AI, cs.CL · submitted Jan 8, 2024 · updated Feb 26, 2024
abstract · pdf · html