about
DeepSeekMath: Pushing Limits of Mathematical Reasoning in Open Language Models (arxiv.org)
1 point by tosh on Jan 28, 2025 | hide | past | pdf | 1 comment on HN

In plain words: A 7-billion-parameter model was further trained on 120 billion math tokens from the web, then tuned by rewarding its own correct answers using less memory than usual. It solved 51.7% of competition-level math problems without tools or voting over several answers, near GPT-4.

Abstract · DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pre-training DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens sourced from Common Crawl, together with natural language and code data. DeepSeekMath 7B has achieved an impressive score of 51.7% on the competition-level MATH benchmark without relying on external toolkits and voting techniques, approaching the performance level of Gemini-Ultra and GPT-4. Self-consistency over 64 samples from DeepSeekMath 7B achieves 60.9% on MATH. The mathematical reasoning capability of DeepSeekMath is attributed to two key factors: First, we harness the significant potential of publicly available web data through a meticulously engineered data selection pipeline. Second, we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.

Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, Daya Guo
arXiv:2402.03300 · cs.CL, cs.AI, cs.LG · submitted Feb 5, 2024 · updated Apr 27, 2024
abstract · pdf · html

add comment on HN
Also discussed: Jan 2025 (6 points, 0 comments) · Feb 2024 (3 points, 0 comments) · Feb 2024 (1 point, 0 comments) · Feb 2024 (2 points, 1 comment) · Feb 2024 (2 points, 0 comments)

if this is what they're releasing to the public, I wonder what they are withholding

but then I remember, they're Chinese... and coming up with a novelty and then making damned sure nobody else can get it is a way to be characteristic of the west... or so I wished were true