about
Every Flop Counts: Scaling a 300B LLM Without Premium GPUs (arxiv.org)
117 points by bretpiatt on Mar 24, 2025 | hide | past | pdf | 9 comments on HN

In plain words: They trained a huge language model that only uses a small slice of its network for each question, running it on cheaper, lower-performance chips with tricks to cut waste and handle training glitches. It matched similar-sized models while cutting initial training computing costs by about 20%.

Abstract · Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs

In this technical report, we tackle the challenges of training large-scale Mixture of Experts (MoE) models, focusing on overcoming cost inefficiency and resource limitations prevalent in such systems. To address these issues, we present two differently sized MoE large language models (LLMs), namely Ling-Lite and Ling-Plus (referred to as "Bailing" in Chinese, spelled Bǎilíng in Pinyin). Ling-Lite contains 16.8 billion parameters with 2.75 billion activated parameters, while Ling-Plus boasts 290 billion parameters with 28.8 billion activated parameters. Both models exhibit comparable performance to leading industry benchmarks. This report offers actionable insights to improve the efficiency and accessibility of AI development in resource-constrained settings, promoting more scalable and sustainable technologies. Specifically, to reduce training costs for large-scale MoE models, we propose innovative methods for (1) optimization of model architecture and training processes, (2) refinement of training anomaly handling, and (3) enhancement of model evaluation efficiency. Additionally, leveraging high-quality data generated from knowledge graphs, our models demonstrate superior capabilities in tool use compared to other models. Ultimately, our experimental findings demonstrate that a 300B MoE LLM can be effectively trained on lower-performance devices while achieving comparable performance to models of a similar scale, including dense and MoE models. Compared to high-performance devices, utilizing a lower-specification hardware system during the pre-training phase demonstrates significant cost savings, reducing computing costs by approximately 20%. The models can be accessed at https://huggingface.co/inclusionAI.

Ling Team, Binwei Zeng, Chao Huang, Chao Zhang, Changxin Tian, Cong Chen, Dingnan Jin, Feng Yu, Feng Zhu, Feng Yuan, Fakang Wang, Gangshan Wang, et al.
arXiv:2503.05139 · cs.LG, cs.AI, cs.CL · submitted Mar 7, 2025 · updated Mar 10, 2025
abstract · pdf · html · 34 pages

add comment on HN
Also discussed: Mar 2025 (2 points, 0 comments)

They never mention what hardware they're on.

Table 1 is the closest thing. Device specs for six devices: 120-989 TFLOPS and 64-96 GB RAM.

An RTX 5090 is about 105 TFLOPS.

https://www.techpowerup.com/gpu-specs/geforce-rtx-5090.c4216

The 96GB (HBM2e) SKU is named PPU from T-head semiconductor (basically a subsidiary of Alibaba). The spec is very similar to H20. Other chips they were using include Huawei Ascend 910B (64GB) and maybe other domestic designed chips.
I was surprised not to see a Kunlun P800 there.
I'm pretty surprised by the claimed memory usage for 300B parameters (table 1). If we compare similar models:

- Llama 3.1 with 405B parameters: 2 TB of memory (FP32), 500 GB (FP8)

- DeepSeek R1 with 671B parameters: 1.3 TB (scaling linearly, around 600 GB for 300B parameters)

Ling claims no more than 96 GB of memory, most likely for inference. That's far more than a 20% reduction. Am I missing something?

I think they only claim their "Ling-Lite" 17B model can fit on a single 96GB GPU, their 300B model needs 8 of them (768GB of HBM)
Some of these models still produce great results with something low like 2.7 bits per variable.
They've shared some interesting optimization techniques for bigger LLMs that's all, not exactly low powered devices as in power consumption. Still a good read.
I think this is the one where they train LLM without NVIDIA GPU's.
They talk about CUDA level tracing in their framework. I assume its just consumer GPU's that Nvidia say arent meant to be used in datacenters.