about
Naturalprover: Grounded Mathematical Proof Generation with Language Models (arxiv.org)
41 points by PaulHoule on May 26, 2022 | hide | past | pdf | 2 comments on HN

In plain words: NaturalProver writes math proofs by leaning on cited theorems and definitions, retrieved or supplied, and can require citations to appear. Math students rated its next-step suggestions and proofs better than a fine-tuned language model, calling suggestions correct and useful over 40% of the time.

Abstract · NaturalProver: Grounded Mathematical Proof Generation with Language Models

Theorem proving in natural mathematical language - the mixture of symbolic and natural language used by humans - plays a central role in mathematical advances and education, and tests aspects of reasoning that are core to intelligence. Yet it has remained underexplored with modern generative models. We study large-scale language models on two new generation tasks: suggesting the next step in a mathematical proof, and full proof generation. We develop NaturalProver, a language model that generates proofs by conditioning on background references (e.g. theorems and definitions that are either retrieved or human-provided), and optionally enforces their presence with constrained decoding. On theorems from the NaturalProofs benchmark, NaturalProver improves the quality of next-step suggestions and generated proofs over fine-tuned GPT-3, according to human evaluations from university-level mathematics students. NaturalProver is capable of proving some theorems that require short (2-6 step) proofs, and providing next-step suggestions that are rated as correct and useful over 40% of the time, which is to our knowledge the first demonstration of these capabilities using neural language models.

Sean Welleck, Jiacheng Liu, Ximing Lu, Hannaneh Hajishirzi, Yejin Choi
arXiv:2205.12910 · cs.CL, cs.AI · submitted May 25, 2022 · updated Oct 31, 2022
abstract · pdf · html · NeurIPS 2022

add comment on HN

I feel like 100 theorems they used for fine-tuning is too small of a dataset to train on.
They should apply “conditional deciding” to programming code generation models. I don’t that’s been done before