about
OpenDiLoCo: Open-Source Framework for Distributed Low-Communication Training (arxiv.org)
4 points by Mougatine on Jul 11, 2024 | hide | past | pdf | discuss on HN

In plain words: Trains one large language model on machines spread around the world, letting each work alone for long stretches before syncing to cut the data exchanged. It trained across continents while keeping 90-95% of compute busy, and scaled up to billion-parameter models.

Abstract · OpenDiLoCo: An Open-Source Framework for Globally Distributed Low-Communication Training

OpenDiLoCo is an open-source implementation and replication of the Distributed Low-Communication (DiLoCo) training method for large language models. We provide a reproducible implementation of the DiLoCo experiments, offering it within a scalable, decentralized training framework using the Hivemind library. We demonstrate its effectiveness by training a model across two continents and three countries, while maintaining 90-95% compute utilization. Additionally, we conduct ablations studies focusing on the algorithm's compute efficiency, scalability in the number of workers and show that its gradients can be all-reduced using FP16 without any performance degradation. Furthermore, we scale OpenDiLoCo to 3x the size of the original work, demonstrating its effectiveness for billion parameter models.

Sami Jaghouar, Jack Min Ong, Johannes Hagemann
arXiv:2407.07852 · cs.LG, cs.DC · submitted Jul 10, 2024
abstract · pdf · html

add comment on HN