In plain words: A set of 371 undergraduate math problems with English statements and proofs plus machine-checkable versions, to test turning math into formal code and proving it. They tried two translation tricks: reusing similar examples as prompts, and turning generated translations back into English.
Abstract
We introduce ProofNet, a benchmark for autoformalization and formal proving of undergraduate-level mathematics. The ProofNet benchmarks consists of 371 examples, each consisting of a formal theorem statement in Lean 3, a natural language theorem statement, and a natural language proof. The problems are primarily drawn from popular undergraduate pure mathematics textbooks and cover topics such as real and complex analysis, linear algebra, abstract algebra, and topology. We intend for ProofNet to be a challenging benchmark that will drive progress in autoformalization and automatic theorem proving. We report baseline results on statement autoformalization via in-context learning. Moreover, we introduce two novel statement autoformalization methods: prompt retrieval and distilled backtranslation.
Zhangir Azerbayev, Bartosz Piotrowski, Hailey Schoelkopf, Edward W. Ayers, Dragomir Radev, Jeremy Avigad
arXiv:2302.12433 · cs.CL, cs.AI, cs.LO · submitted Feb 24, 2023
abstract · pdf · html