In plain words: A text-based AI is fed the predictions a vision model makes on a batch of medical images and judges them together, flagging wrong answers and correcting them without retraining. Tests showed it caught errors and raised the vision model's accuracy.
Abstract · GPT4MIA: Utilizing Generative Pre-trained Transformer (GPT-3) as A Plug-and-Play Transductive Model for Medical Image Analysis
In this paper, we propose a novel approach (called GPT4MIA) that utilizes Generative Pre-trained Transformer (GPT) as a plug-and-play transductive inference tool for medical image analysis (MIA). We provide theoretical analysis on why a large pre-trained language model such as GPT-3 can be used as a plug-and-play transductive inference model for MIA. At the methodological level, we develop several technical treatments to improve the efficiency and effectiveness of GPT4MIA, including better prompt structure design, sample selection, and prompt ordering of representative samples/features. We present two concrete use cases (with workflow) of GPT4MIA: (1) detecting prediction errors and (2) improving prediction accuracy, working in conjecture with well-established vision-based models for image classification (e.g., ResNet). Experiments validate that our proposed method is effective for these two tasks. We further discuss the opportunities and challenges in utilizing Transformer-based large language models for broader MIA applications.
Yizhe Zhang, Danny Z. Chen
arXiv:2302.08722 · cs.CV, cs.AI, cs.LG, cs.MM · submitted Feb 17, 2023 · updated Mar 21, 2023
abstract · pdf · html · Version 3: Added appendix with more results and visualizations. Questions and suggestions are welcome