In plain words: A collection of 12,504 formal specifications for testing whether AI can write code that is mathematically proven correct, not code that looks right. Off-the-shelf AI solved 82% of the Dafny tasks but far fewer in the other two languages; plain-English hints did not help.
Abstract
We present and test the largest benchmark for vericoding, LLM-generation of formally verified code from formal specifications - in contrast to vibe coding, which generates potentially buggy code from a natural language description. Our benchmark contains 12,504 formal specifications, with 3,029 in Dafny, 2,334 in Verus/Rust and 7,141 in Lean. Of these, 6,174 are new unseen problems. We find vericoding success rates of 27% in Lean, 44% in Verus/Rust and 82% in Dafny using off-the-shelf LLMs. Adding natural-language descriptions does not significantly improve performance. We also find that LLM progress has improved progress on pure Dafny verification from 68% to 96% over the past year. The benchmark and vericoding results are shared at https://github.com/Beneficial-AI-Foundation/vericoding-benchmark
Sergiu Bursuc, Theodore Ehrenborg, Shaowei Lin, Lacramioara Astefanoaei, Ionel Emilian Chiosa, Jure Kukovec, Alok Singh, Oliver Butterley, Adem Bizid, Quinn Dougherty, Miranda Zhao, Max Tan, et al.
arXiv:2509.22908 · cs.SE, cs.LG, cs.PL · submitted Sep 26, 2025
abstract · pdf · html · 25 pages, 1 figure; data available at https://github.com/Beneficial-AI-Foundation/vericoding-benchmark