about
Hecaton: Training and Finetuning LLMs with Scalable Chiplet Systems (arxiv.org)
2 points by PaulHoule on Sep 12, 2024 | hide | past | pdf | discuss on HN

In plain words: Instead of one giant chip, it splits training across small chips packed together with fast links, with scheduling that cuts memory traffic and on-chip chatter. On a huge language model, it ran 5.29 times faster than the usual way of splitting training across chips.

Abstract · Hecaton: Training Large Language Models with Scalable Chiplet Systems

Large Language Models (LLMs) have achieved remarkable success in various fields, but their training and finetuning require massive computation and memory, necessitating parallelism which introduces heavy communication overheads. Driven by advances in packaging, the chiplet architecture emerges as a potential solution, as it can integrate computing power, as well as utilize on-package links with better signal integrity, higher bandwidth, and lower energy consumption. However, most existing chiplet-related works focus on DNN inference. Directly porting them to LLM training introduces significantly large quantities of DRAM access and network-on-package (NoP) overheads which make state-of-the-art chiplet designs fail, highlighting a research gap. This work proposes Hecaton, a scalable and cost-effective chiplet system for LLM training. We first provide a chiplet architecture with tailored scheduling that can largely reduce DRAM accesses. We further design an efficient distributed training method that reduces NoP communication complexity and relieves constraints on SRAM capacity and layout. Theoretical analysis shows that the entire system achieves weak scaling: as the workload and hardware resources grow proportionally, the computation-to-communication ratio remains nearly constant. Experiments with various workloads and hardware configurations verify the property, and Hecaton achieves $5.29\times$ performance improvement and $3.46\times$ energy reduction on Llama3.1-405B, compared to the tensor parallelism in Megatron. To the best of our knowledge, we propose the first chiplet architecture specifically used for LLM training or finetuning, with guaranteed performance regardless of the problem scale.

Zongle Huang, Shupei Fan, Chen Tang, Xinyuan Lin, Shuwen Deng, Yongpan Liu
arXiv:2407.05784 · cs.AR · submitted Jul 8, 2024 · updated Nov 27, 2024
abstract · pdf · html

add comment on HN