In plain words: A language model trained first on general text and then on technical writing learns to solve math and science problems with no calculator or equation solver. It beat earlier models on technical benchmarks and answered nearly a third of over 200 college-level science questions.
Abstract
Language models have achieved remarkable performance on a wide range of tasks that require natural language understanding. Nevertheless, state-of-the-art models have generally struggled with tasks that require quantitative reasoning, such as solving mathematics, science, and engineering problems at the college level. To help close this gap, we introduce Minerva, a large language model pretrained on general natural language data and further trained on technical content. The model achieves state-of-the-art performance on technical benchmarks without the use of external tools. We also evaluate our model on over two hundred undergraduate-level problems in physics, biology, chemistry, economics, and other sciences that require quantitative reasoning, and find that the model can correctly answer nearly a third of them.
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, et al.
arXiv:2206.14858 · cs.CL, cs.AI, cs.LG · submitted Jun 29, 2022 · updated Jul 1, 2022
abstract · pdf · html · 12 pages, 5 figures + references and appendices