about
Demo: Decoupled Momentum Optimization – training neural networks in parallel (arxiv.org)
3 points by sergiotapia on Dec 2, 2024 | hide | past | pdf | discuss on HN

In plain words: Instead of sending every gradient value between GPUs each step, each GPU keeps its own momentum and sends only the largest components after a reshaping transform, carrying leftovers forward. This cut data sent per GPU by up to 85x with similar loss and accuracy.

Abstract · DeMo: Decoupled Momentum Optimization

Scaling neural network training increasingly depends on synchronous data-parallelism, yet full-precision gradient all-reduce imposes a severe communication bottleneck. We propose Decoupled Momentum Optimization (DeMo), a drop-in replacement for any momentum-based optimizers that significantly reduces the communication bandwidth while maintaining convergence. DeMo (i) decouples local momentum updates, (ii) applies a fast orthonormal transform (e.g., DCT) followed by top-k sparsification, and (iii) reuses the momentum buffer as error feedback via momentum subtraction. This design reduces per-step communication by up to two orders of magnitude with minimal computational overhead. Experiments on 300M and 1B-parameter DeMo language models show DeMo transmits up to 85x less data per GPU than AdamW-DDP while achieving comparable loss and accuracy. DeMo is topology-agnostic and enables training across multi-datacenter or Ethernet-based setups. Code is available at https://github.com/bloc97/DeMo

Bowen Peng, Lizhang Chen, Baiyu Su, Jeffrey Quesnelle, Diederik P. Kingma, Qiang Liu
arXiv:2411.19870 · cs.LG, cs.AI · submitted Nov 29, 2024 · updated Feb 6, 2026
abstract · pdf · html

add comment on HN
Also discussed: Dec 2024 (1 point, 0 comments)