about
MathPrompter: Mathematical Reasoning Using LLMs. New SOTA on MultiArith(92.5%) (arxiv.org)
1 point by famouswaffles on Mar 13, 2023 | hide | past | pdf | 1 comment on HN

In plain words: It turns each math problem into several algebraic expressions or Python functions that solve it in different ways, then checks whether they agree before answering. On a standard arithmetic problem set, this raised accuracy from 78.7% to 92.5% over the best earlier prompt-based approach.

Abstract · MathPrompter: Mathematical Reasoning using Large Language Models

Large Language Models (LLMs) have limited performance when solving arithmetic reasoning tasks and often provide incorrect answers. Unlike natural language understanding, math problems typically have a single correct answer, making the task of generating accurate solutions more challenging for LLMs. To the best of our knowledge, we are not aware of any LLMs that indicate their level of confidence in their responses which fuels a trust deficit in these models impeding their adoption. To address this deficiency, we propose `MathPrompter', a technique that improves performance of LLMs on arithmetic problems along with increased reliance in the predictions. MathPrompter uses the Zero-shot chain-of-thought prompting technique to generate multiple Algebraic expressions or Python functions to solve the same math problem in different ways and thereby raise the confidence level in the output results. This is in contrast to other prompt based CoT methods, where there is no check on the validity of the intermediate steps followed. Our technique improves over state-of-the-art on the MultiArith dataset ($78.7\%\rightarrow92.5\%$) evaluated using 175B parameter GPT-based LLM.

Shima Imani, Liang Du, Harsh Shrivastava
arXiv:2303.05398 · cs.CL, cs.AI · submitted Mar 4, 2023
abstract · pdf · html

add comment on HN
Also discussed: Mar 2023 (1 point, 0 comments)

Large Language Models (LLMs) have limited performance when solving arithmetic reasoning tasks and often provide incorrect answers. Unlike natural language understanding, math problems typically have a single correct answer, making the task of generating accurate solutions more challenging for LLMs. To the best of our knowledge, we are not aware of any LLMs that indicate their level of confidence in their responses which fuels a trust deficit in these models impeding their adoption. To address this deficiency, we propose `MathPrompter', a technique that improves performance of LLMs on arithmetic problems along with increased reliance in the predictions. MathPrompter uses the Zero-shot chain-of-thought prompting technique to generate multiple Algebraic expressions or Python functions to solve the same math problem in different ways and thereby raise the confidence level in the output results. This is in contrast to other prompt based CoT methods, where there is no check on the validity of the intermediate steps followed. Our technique improves over state-of-the-art on the MultiArith dataset (78.7%→92.5%) evaluated using 175B parameter GPT-based LLM.