about
UniPool: A Globally Shared Expert Pool for Mixture-of-Experts (arxiv.org)
2 points by danborn26 148 days ago | hide | past | pdf | discuss on HN

In plain words: Instead of fixing each layer's experts from the start, layers briefly share one pool to learn which experts fit where, then lock that split and train as usual. At equal compute, held-out loss fell 0.024–0.037 versus the usual fixed setup, barely changing speed.

Abstract · UniPool: Learning Expert-to-Layer Ownership from Brief Global Access

Most Mixture-of-Experts (MoE) transformers preassign each expert to one layer for the whole of training. We study whether a brief global-access phase can learn a compatible expert-to-layer allocation that is then executed privately. UniPool-Lock gives every layer its own router over one global pool, balances aggregate pool usage, and scores experts with a scale-stable NormRouter. After the short ownership-learning phase, it assigns each expert to one layer, locks the disjoint allocation, and continues training as a layer-private MoE. With the same expert-FFN budget and routed expert FLOPs, this recipe lowers held-out loss relative to vanilla MoE by 0.024-0.037 across five dense-equivalent scales from 182M to 1.5B, including -0.0247 at 1.5B after 60B tokens. The ownership-learning phase lasts 2K steps, about 3.3% of training, and post-lock step time is within -0.7% to +2.8% of vanilla MoE in our throughput measurements. Controls locate the gain in full-pool training and in the allocation it produces: a random disjoint allocation fixed at initialization, trained with the same router and losses, matches vanilla, whereas locking the allocation learned during the full-pool phase recovers nearly all of the persistent full-pool gain, and substituting a random allocation at the lock forfeits about half of it. Keeping full-pool access throughout training (UniPool-Full) also improves over vanilla MoE at four scales and outperforms it with only 66.7% (182M) to 50% (469M and 650M) of its expert parameters.

Minbin Huang, Han Shi, Chuanyang Zheng, Yimeng Wu, Guoxuan Chen, Xintong Yu, Yichun Yin, Hong Cheng
arXiv:2605.06665 · cs.LG, cs.AI · submitted May 7, 2026 · updated Sep 27, 2026
abstract · pdf · html

add comment on HN