In plain words: Board-certified doctors and residents tested the image-reading AI on CT scans, MRIs, ECGs, and clinical photos to see if it could diagnose conditions. It could describe what it saw, but its diagnoses and treatment decisions were poor enough to risk patient safety.
Abstract · GPT-4V(ision) Unsuitable for Clinical Care and Education: A Clinician-Evaluated Assessment
OpenAI's large multimodal model, GPT-4V(ision), was recently developed for general image interpretation. However, less is known about its capabilities with medical image interpretation and diagnosis. Board-certified physicians and senior residents assessed GPT-4V's proficiency across a range of medical conditions using imaging modalities such as CT scans, MRIs, ECGs, and clinical photographs. Although GPT-4V is able to identify and explain medical images, its diagnostic accuracy and clinical decision-making abilities are poor, posing risks to patient safety. Despite the potential that large language models may have in enhancing medical education and delivery, the current limitations of GPT-4V in interpreting medical images reinforces the importance of appropriate caution when using it for clinical decision-making.
Senthujan Senkaiahliyan, Augustin Toma, Jun Ma, An-Wen Chan, Andrew Ha, Kevin R. An, Hrishikesh Suresh, Barry Rubin, Bo Wang
arXiv:2403.12046 · cs.CV · submitted Nov 14, 2023
abstract · pdf · html
> AI company releases generalist model for testing/experimentation
> Users unwisely treat it like a universal oracle and give it tasks far outside its training domain
> It doesn't perform well
> People are shocked and warn about the "dangers of AI"
This happens every time. Why can't we treat AI tools like they actually are: interesting demonstrations of emergent intelligent properties that are a few versions away from production-ready capabilities?