about
MuMath-Code on ArXiv (arxiv.org)
2 points by pajop on Jul 8, 2024 | hide | past | pdf | discuss on HN

In plain words: They turned old math problems into new ones, wrote solutions mixing reasoning with Python code, and trained an open model in two steps—text first, then code—so it can run code in Python. The 70B version scored 90.7% on GSM8K, the best among open models.

Abstract · MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning

The tool-use Large Language Models (LLMs) that integrate with external Python interpreters have significantly enhanced mathematical reasoning capabilities for open-source LLMs, while tool-free methods chose another track: augmenting math reasoning data. However, a great method to integrate the above two research paths and combine their advantages remains to be explored. In this work, we firstly include new math questions via multi-perspective data augmenting methods and then synthesize code-nested solutions to them. The open LLMs (i.e., Llama-2) are finetuned on the augmented dataset to get the resulting models, MuMath-Code ($μ$-Math-Code). During the inference phase, our MuMath-Code generates code and interacts with the external python interpreter to get the execution results. Therefore, MuMath-Code leverages the advantages of both the external tool and data augmentation. To fully leverage the advantages of our augmented data, we propose a two-stage training strategy: In Stage-1, we finetune Llama-2 on pure CoT data to get an intermediate model, which then is trained on the code-nested data in Stage-2 to get the resulting MuMath-Code. Our MuMath-Code-7B achieves 83.8 on GSM8K and 52.4 on MATH, while MuMath-Code-70B model achieves new state-of-the-art performance among open methods -- achieving 90.7% on GSM8K and 55.1% on MATH. Extensive experiments validate the combination of tool use and data augmentation, as well as our two-stage training strategy. We release the proposed dataset along with the associated code for public use.

Shuo Yin, Weihao You, Zhilong Ji, Guoqiang Zhong, Jinfeng Bai
arXiv:2405.07551 · cs.CL, cs.AI · submitted May 13, 2024
abstract · pdf · html · The state-of-the-art open-source tool-use LLMs for mathematical reasoning

add comment on HN