In plain words: A tool that lists possible diseases from a patient's history and tests was tested alone and as a doctor's helper on 302 hard cases. It beat unassisted clinicians, and doctors using it got the right diagnosis in their top 10 51.7% of the time, beating search alone.
Abstract · Towards Accurate Differential Diagnosis with Large Language Models
An accurate differential diagnosis (DDx) is a cornerstone of medical care, often reached through an iterative process of interpretation that combines clinical history, physical examination, investigations and procedures. Interactive interfaces powered by Large Language Models (LLMs) present new opportunities to both assist and automate aspects of this process. In this study, we introduce an LLM optimized for diagnostic reasoning, and evaluate its ability to generate a DDx alone or as an aid to clinicians. 20 clinicians evaluated 302 challenging, real-world medical cases sourced from the New England Journal of Medicine (NEJM) case reports. Each case report was read by two clinicians, who were randomized to one of two assistive conditions: either assistance from search engines and standard medical resources, or LLM assistance in addition to these tools. All clinicians provided a baseline, unassisted DDx prior to using the respective assistive tools. Our LLM for DDx exhibited standalone performance that exceeded that of unassisted clinicians (top-10 accuracy 59.1% vs 33.6%, [p = 0.04]). Comparing the two assisted study arms, the DDx quality score was higher for clinicians assisted by our LLM (top-10 accuracy 51.7%) compared to clinicians without its assistance (36.1%) (McNemar's Test: 45.7, p < 0.01) and clinicians with search (44.4%) (4.75, p = 0.03). Further, clinicians assisted by our LLM arrived at more comprehensive differential lists than those without its assistance. Our study suggests that our LLM for DDx has potential to improve clinicians' diagnostic reasoning and accuracy in challenging cases, meriting further real-world evaluation for its ability to empower physicians and widen patients' access to specialist-level expertise.
Daniel McDuff, Mike Schaekermann, Tao Tu, Anil Palepu, Amy Wang, Jake Garrison, Karan Singhal, Yash Sharma, Shekoofeh Azizi, Kavita Kulkarni, Le Hou, Yong Cheng, et al.
arXiv:2312.00164 · cs.CY, cs.AI · submitted Nov 30, 2023
abstract · pdf · html
I have a friend with Crohn's who was feeling low energy. I was a gym bro at the time and convinced him to take a testosterone test (because all problems re caused by low T when you're a gym bro).
His doctor wouldn't even entertain the idea, saying he's a young man and it's very unlikely that he'd have low T. He did the test privately and his T is significantly below normal. If you Google, there are actually many papers showing correlation between Crohn's and low T. I bet an AI would find it.
Similarly, doctors missed my mum's recent cancer diagnosis. She also had factors that would make her more susceptible to breast cancer, googling finds many papers that show causation.
The problem is that those things aren't extremely common and haven't made their way to NICE guidelines or whatever GPs use.
Not that I'm blaming doctors, they have 10 minute appointments and don't have time to do anything. I'm sure AI would recommend significantly more lab tests which would put even more pressure on the NHS.