about
LLMs exceed physicians on complex text-based differential diagnosis (arxiv.org)
3 points by rippeltippel 233 days ago | hide | past | pdf | 2 comments on HN

In plain words: CaBot turns a short case description into a full written and narrated diagnostic presentation, reasoning step by step like an expert doctor instead of just guessing the final answer. In blind tests, physicians mistook its differentials for human-written ones in 74% of trials.

Abstract · Teaching large language models to reason like expert diagnosticians

Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferences (CPCs), published continuously since 1923, feature expert physicians who demonstrate diagnostic reasoning to peers, and have been used for decades to evaluate AI. However, prior AI evaluations have largely focused on final diagnostic accuracy rather than nuanced clinical reasoning. Here, we introduce Dr. CaBot, an agentic AI system that emulates an expert diagnostician by generating written and narrated slide-based presentations from an initial case description alone. CaBot recently generated the first AI diagnosis published in the 100+ year history of the NEJM CPCs. In blinded evaluations, physicians misclassified the source of the differential (CaBot vs. physician-written) in 46/62 (74%) of trials and rated them favorably across quality dimensions. When tasked with solving cases for 72 patients with undiagnosed disease from the NIH Undiagnosed Diseases Network, CaBot identified the working diagnosis in 50/72 (69%) of cases from referral notes alone. To promote transparency and research, we also developed CPC-Bench, a physician-validated benchmark based on 7,102 CPCs and 47,648 questions across 10 tasks. We show that CaBot outperforms frontier models on CPC-Bench, and release both CaBot and CPC-Bench publicly to foster progress in clinical AI.

Thomas A. Buckley, Riccardo Conci, Peter G. Brodeur, Jason Gusdorf, Sourik Beltrán, Bita Behrouzi, Byron Crowe, Jacob Dockterman, Muzzammil Muhammad, Sarah Ohnigian, Andrew Sanchez, James A. Diao, et al.
arXiv:2509.12194 · cs.AI, cs.CV · submitted Sep 15, 2025 · updated May 24, 2026
abstract · pdf

add comment on HN
Also discussed: Oct 2025 (3 points, 1 comment)

This was using o3. GPT 5.2/5.3 should be much improved.

Just like software engineering, it may be best to leave it up to the AI to do the work but let a human guide it and check it.

I wonder if we’ll have to develop strategies for battling confirmation bias. Human review only works if the review is independent.