about
CodeLogician: Neuro-symbolic reasoning for precise software analysis (arxiv.org)
2 points by NTCTech 230 days ago | hide | past | pdf | 1 comment on HN

In plain words: An agent has a language model write a formal math description of a program, then a proof engine answers questions about its states, paths, and edge cases. On a new test, it closed a 41-47 percentage point accuracy gap over language models working alone.

Abstract · Imandra CodeLogician: Neuro-Symbolic Reasoning for Precise Analysis of Software Logic

Large Language Models (LLMs) have shown strong performance on code understanding tasks, yet they fundamentally lack the ability to perform precise, exhaustive mathematical reasoning about program behavior. Existing benchmarks either focus on mathematical proof automation, largely disconnected from real-world software, or on engineering tasks that do not require semantic rigor. We present CodeLogician, a neurosymbolic agent for precise analysis of software logic, integrated with ImandraX, an industrial automated reasoning engine deployed in financial markets and safety-critical systems. Unlike prior approaches that use formal methods primarily to validate LLM outputs, CodeLogician uses LLMs to construct explicit formal models of software systems, enabling automated reasoning to answer rich semantic questions beyond binary verification outcomes. To rigorously evaluate mathematical reasoning about software logic, we introduce code-logic-bench, a benchmark targeting the middle ground between theorem proving and software engineering benchmarks. It measures reasoning correctness about program state spaces, control flow, coverage constraints, and edge cases, with ground truth defined via formal modeling and region decomposition. Comparing LLM-only reasoning against LLMs augmented with CodeLogician, formal augmentation yields substantial improvements, closing a 41-47 percentage point gap in reasoning accuracy. These results demonstrate that neurosymbolic integration is essential for scaling program analysis toward rigorous, autonomous software understanding.

Hongyu Lin, Samer Abdallah, Makar Valentinov, Paul Brennan, Elijah Kagan, Christoph M. Wintersteiger, Denis Ignatovich, Grant Passmore
arXiv:2601.11840 · cs.AI, cs.LO, cs.SE · submitted Jan 17, 2026 · updated Feb 6, 2026
abstract · pdf · html · 52 pages, 23 figures. Includes a new benchmark dataset (code-logic-bench) and evaluation of neurosymbolic reasoning for software analysis

add comment on HN

Found this via a related paper on Lobste.rs today. The author makes a compelling argument that we've hit the limit of "Vibe Coding" (LLMs guessing via tokens) and need to move to "State-Space Exploration" (Formal Verification).

They claim their neuro-symbolic approach closes a 40%+ accuracy gap in reasoning tasks by forcing the LLM to construct a formal model rather than just predicting the next token. Curious if anyone has tried their CodeLogician agent yet, or if this is just more symbolic AI hype?