In plain words: It is a language model that keeps 671 billion numbers but uses only 37 billion per word, spreading work across expert parts and balancing them without extra penalty terms. It beat other open models and matched the best closed ones, training cheaply, never crashing.
Abstract
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning and Reinforcement Learning stages to fully harness its capabilities. Comprehensive evaluations reveal that DeepSeek-V3 outperforms other open-source models and achieves performance comparable to leading closed-source models. Despite its excellent performance, DeepSeek-V3 requires only 2.788M H800 GPU hours for its full training. In addition, its training process is remarkably stable. Throughout the entire training process, we did not experience any irrecoverable loss spikes or perform any rollbacks. The model checkpoints are available at https://github.com/deepseek-ai/DeepSeek-V3.
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, et al.
arXiv:2412.19437 · cs.CL, cs.AI · submitted Dec 27, 2024 · updated Feb 18, 2025
abstract · pdf · html
2,788,000 GPU-hours * 350W TDP of H800 = 975,800,000 GPU Watt-hours
975,800,000 GPU Wh * (1.2 to account for non-GPU hardware) * (1.3 PUE [1]) = 1,522,248,000 Total Wh, or 1,522,248 kWh to train DeepSeek-V3
(1,522,248 kWh) * (0.582kg CO2eq/kWh in China [2]) = 885,948 kg CO2 equivalents to train DeepSeek-V3
A typical US passenger vehicle emits about 4.6 metric tons of CO2 per year. [3]
885,948 kg CO2 per DeepSeek / 4,600 kg CO2 per car = 192.6 cars per DeepSeek
So, the final training run for DeepSeek-V3 emitted as much greenhouse gasses as would be emitted from running about 193 more cars on the road for a year.
I also did some more math and found that this training run used about as much electricity as 141 US households would use over the course of a year. [4]
[1] https://enviliance.com/regions/east-asia/cn/report_10060
[2] https://ourworldindata.org/grapher/carbon-intensity-electric...
[3] https://www.epa.gov/greenvehicles/greenhouse-gas-emissions-t...
[4] divided total kWh by the value here: https://www.eia.gov/tools/faqs/faq.php?id=97&t=3