about
UBEP: Expert Parallelism Communication Library for Production Superpods (arxiv.org)
1 point by Jimmc414 86 days ago | hide | past | pdf | discuss on HN

In plain words: UBEP is a communication library for Mixture-of-Experts models on big superpods that runs its data-sending steps at once instead of one after another and cuts waiting between them. It cut communication delay by up to 52.4% and time per generated word by up to 11.1%.

Abstract · UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods

The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatrix384, introduces critical challenges beyond raw interconnect bandwidth. While these systems provide unified global address spaces and high-bandwidth fabrics, their full potential for sparse MoE communication is hindered by three fundamental bottlenecks: (1) Strict execution serialization imposed by coarse-grained Bulk Synchronous Parallel (BSP) orchestration of interdependent communication phases; (2) Prohibitive synchronization overhead that fails to scale alongside high interconnect bandwidth; and (3) Severe load imbalance resulting from distance-agnostic scheduling of irregular token traffic. To eliminate these bottlenecks, we introduce UBEP (Unified-Bus Expert Parallelism), a production-ready communication library that rethinks MoE's All-to-All primitives for modern superpod architectures. Through large scale experiments, UBEP reduces All-to-All latency by up to 52.4% and MoE inference Time Per Output Token (TPOT) by up to 11.1%.

Yipeng Liu, Chang Liu, Si Shen, Jiaqi Zheng, Mingfan Li, Yuyang Yang, Guanhua Li, Yuquan Zhang, Yimeng Xu, Zhongzhe Hu, Zhiyuan Huang, Qihang Duan, et al.
arXiv:2607.06202 · cs.DC, cs.AI, cs.NI · submitted Jul 7, 2026 · updated Jul 8, 2026
abstract · pdf · html · Accepted to ACM SIGCOMM 2026. Corresponding authors: [email protected] (J. Zheng), [email protected] (Z. Hu)

add comment on HN