about
Toward Guarantees for Clinical Reasoning in Vision Language Models (arxiv.org)
5 points by barthelomew 216 days ago | hide | past | pdf | 3 comments on HN

In plain words: A checker turns a model's written X-ray findings into formal logic, then uses a solver to judge each diagnosis as supported, made up, or missing. Across seven models and five X-ray sets, it caught errors wording scores miss, and blocking unsupported claims removed hallucinations.

Abstract · Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification

Vision-language models (VLMs) show promise in drafting radiology reports, yet they frequently suffer from logical inconsistencies, generating diagnostic impressions unsupported by their own perceptual findings or missing logically entailed conclusions. Standard lexical metrics heavily penalize clinical paraphrasing and fail to capture these deductive failures in reference-free settings. Toward guarantees for clinical reasoning, we introduce a neurosymbolic verification framework that deterministically audits the internal consistency of VLM-generated reports. Our pipeline autoformalizes free-text radiographic findings into structured propositional evidence, utilizing an SMT solver (Z3) and a clinical knowledge base to verify whether each diagnostic claim is mathematically entailed, hallucinated, or omitted. Evaluating seven VLMs across five chest X-ray benchmarks, our verifier exposes distinct reasoning failure modes, such as conservative observation and stochastic hallucination, that remain invisible to traditional metrics. On labeled datasets, enforcing solver-backed entailment acts as a rigorous post-hoc guarantee, systematically eliminating unsupported hallucinations to significantly increase diagnostic soundness and precision in generative clinical assistants.

Vikash Singh, Debargha Ganguly, Haotian Yu, Chengwei Zhou, Prerna Singh, Brandon Lee, Vipin Chaudhary, Gourav Datta
arXiv:2602.24111 · cs.CV, cs.AI, cs.CL, cs.LO · submitted Feb 27, 2026
abstract · pdf · html

add comment on HN
Also discussed: Mar 2026 (2 points, 0 comments)

AI (VLM-based) radiology models can sound confident and still be wrong ; hallucinating diagnoses that their own findings don't support. This is a silent, and dangerous failure mode.

Our new paper introduces a verification layer that checks every diagnostic claim an AI makes before it reaches a clinician. When our system says a diagnosis is supported, it's been mathematically proven - not just guessed. Every model we tested improved significantly after verification, with our best result hitting 99% soundness.

We're excited about what comes next in building verifiably correct AI systems.

nice work!I know about your work, similar to this: https://arxiv.org/abs/2601.20055 and https://github.com/DebarghaG/proofofthought
Yes, indeed! This work uses the Proof of Thought library and several techniques from VERGE!