Diagnostic Accuracy of ChatGPT Plus (GPT-4o) for the Interpretation of Chest and Extremity Radiographs Against Routine Radiologist Reporting: A Single-Centre, Retrospective, Cross-Sectional Study
Aug 2026· Indian Journal of Radiology and Imaging· 0 citations· 28 references
TL;DR
ChatGPT Plus achieved overall diagnostic accuracy close to routine radiologist reports, with substantial agreement and high specificity but lower sensitivity, particularly in extremity radiographs.
Abstract
Background Multimodal large language models are now widely accessible, but their diagnostic capability on plain-film radiographs is poorly characterized. Most evaluations in radiology address purpose-built convolutional networks rather than general-purpose conversational assistants. Materials and Methods A single-center, retrospective, cross-sectional diagnostic-accuracy study was conducted over 6 months at a tertiary-care teaching hospital in western India. We randomly drew 385 chest and extremity radiographs from PACS, each interpreted independently by ChatGPT Plus (GPT-4o) and compared with the verified radiologist report as the reference standard. Outcomes were sensitivity, specificity, accuracy, likelihood ratios, the diagnostic odds ratio (DOR), Cohen's kappa, and the McNemar exact test, with prespecified subgroup analyses and Wilson 95% confidence intervals. Blinded adjudication of the 39 discordant pairs by an independent consultant radiologist was performed as a sensitivity analysis of the reference standard. Results Among 385 radiographs (228 chest, 157 extremity; abnormal prevalence 25.7%), ChatGPT Plus achieved a sensitivity of 78.8% (95% CI 69.7–85.7), specificity of 93.7% (90.3–96.0), accuracy of 89.9% (86.5–92.5), positive likelihood ratio of 12.5, and a DOR of 55.3. Inter-rater agreement was substantial ( κ = 0.73; 0.65–0.81), with no systematic discordance (McNemar exact p = 0.749). Sensitivity was higher for chest than for extremity radiographs (87.3 vs. 63.9%; Fisher's exact p = 0.010); specificities were comparable. Independent adjudication of the 39 discordant pairs reclassified [X/21] false negatives and [Y/18] false positives as confirmed model errors, [A] as confirmed radiologist omissions or borderline calls, and [B] as legitimately equivocal; the corrected-reference-standard sensitivity and specificity were [S%] and [Sp%] respectively. Conclusion ChatGPT Plus achieved overall diagnostic accuracy close to routine radiologist reports, with substantial agreement and high specificity but lower sensitivity, particularly in extremity radiographs. The model is a plausible supervised educational adjunct or alert application rather than a substitute for expert interpretation; local validation and a human-in-the-loop pathway are prerequisites for any clinical role.
Multimodal large language models can interpret medical images, but their performance for pediatric elbow radiographs remains uncertain. We evaluated the diagnostic performance of GPT-5.2 Instant as accessed through the ChatGPT web interface during the defined study period.
In this prospective, single-cente...
O. Taş, Mehmet Yorgun, R. Aktaş et al.· BMC Medical Imaging· 0 citations
Background/Objectives: This study aimed to compare the diagnostic performance of multimodal (image-capable) large language models (LLMs) and oral and maxillofacial radiologists in detecting nine predefined incidental findings on panoramic radiographs, using a cone-beam computed tomography (CBCT)-derived reference stand...
İsmail Çapar, Utku Cem Hasırcı, Didem Dumanlı Kusay et al.· Healthcare· 0 citations
Accurate interpretation of spine imaging is essential for clinical decision-making, yet the diagnostic potential of large language models (LLMs) for radiological report analysis remains inadequately evaluated in terms of sample size, multi-model comparison, reproducibility, and cross-institutional generalisability. Her...
Hao-Lai Liu, Hao Zhang, Hai-Xin Wei et al.· npj Digital Medicine· 0 citations
Background/Objective: Large language models (LLMs) have shown exam-level performance, yet their reliability and safety in laboratory medicine—where quantitative data interpretation is central—remain insufficiently validated. This study compared the accuracy, interpretive quality, and safety of ChatGPT-5.2, Gemini 3 Pro...
Chest radiography (CXR) is the conventional imaging standard for diagnosing community-acquired pneumonia (CAP) in children. However, concerns regarding diagnostic variability and ionizing radiation exposure have prompted interest in alternative modalities. Point-of-care ultrasound (POCUS) has emerged as a promising, ra...
A. Shaban, Hany A. Zaki, E. Shaban et al.· Egyptian Pediatric Associati...· 0 citations
We developed a national questionnaire to assess radiology reports (RR) physicians’ reading habits and identify their expectations towards RR, focusing on CT and MRI reports. An anonymized 20-item questionnaire was nationally distributed in France via practice and hospital mailing lists and one social media group, and i...
Yasmine Kassab, Matthieu Bailly, Sixtine Brabant et al.· Insights into Imaging· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.