about
Solving Moe Load Imbalance in LLM Training via Optimal Transport (arxiv.org)
1 point by nullnonenilNULL 47 days ago | hide | past | pdf | discuss on HN

In plain words: When expert units get overloaded, this system copies them onto idle machines, choosing where each copy goes by how far its weights travel across the network. Unlike methods that only chase load balance, it cut communication cost up to 74% and sped training 1.43 times.

Abstract · TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training

Mixture-of-Experts (MoE) has become a key architecture for scaling large language models (LLMs), yet its dynamic routing causes severe load imbalance in expert-parallel training. Existing dynamic-replica methods copy hot experts onto idle ranks to share computation, but they optimize load balance alone and ignore the cost of moving expert weights across a multi-node topology, so the resulting cross-node communication can outweigh the balancing gain and inflate training cost. We present TAOT, a topology-aware optimal transport method for dynamic expert-replica placement. TAOT models the overload on hot ranks and the spare capacity on lightly loaded ranks as a balanced entropy-regularized optimal transport problem with a communication-cost matrix, solves it with Sinkhorn-Knopp iterations to produce rank-level flow hints, and combines integer replica matching with token assignment into an executable schedule. At the system level, it overlaps guest-weight transfer with home-expert computation to hide the communication overhead. Experiments show TAOT achieves a 1.43x end-to-end MoE training speedup, reaches balance quality competitive with or better than existing state-of-the-art methods, and attains the lowest weighted expert-communication cost across all configurations, with up to a 74% reduction.

Lingyun Zhang, Henghua Zhang, Shilei Gu, Kai Mo, Shuai Han, Shiyong Li, Yanpeng Wang, Dou Shen
arXiv:2608.03676 · cs.DC · submitted Aug 4, 2026
abstract · pdf · html

add comment on HN