Skip to content
Open access

Diagnostic Accuracy of ChatGPT Plus (GPT-4o) for the Interpretation of Chest and Extremity Radiographs Against Routine Radiologist Reporting: A Single-Centre, Retrospective, Cross-Sectional Study

Aug 2026 · Indian Journal of Radiology and Imaging · 0 citations · 28 references

TL;DR

ChatGPT Plus achieved overall diagnostic accuracy close to routine radiologist reports, with substantial agreement and high specificity but lower sensitivity, particularly in extremity radiographs.

Abstract

Background Multimodal large language models are now widely accessible, but their diagnostic capability on plain-film radiographs is poorly characterized. Most evaluations in radiology address purpose-built convolutional networks rather than general-purpose conversational assistants. Materials and Methods A single-center, retrospective, cross-sectional diagnostic-accuracy study was conducted over 6 months at a tertiary-care teaching hospital in western India. We randomly drew 385 chest and extremity radiographs from PACS, each interpreted independently by ChatGPT Plus (GPT-4o) and compared with the verified radiologist report as the reference standard. Outcomes were sensitivity, specificity, accuracy, likelihood ratios, the diagnostic odds ratio (DOR), Cohen's kappa, and the McNemar exact test, with prespecified subgroup analyses and Wilson 95% confidence intervals. Blinded adjudication of the 39 discordant pairs by an independent consultant radiologist was performed as a sensitivity analysis of the reference standard. Results Among 385 radiographs (228 chest, 157 extremity; abnormal prevalence 25.7%), ChatGPT Plus achieved a sensitivity of 78.8% (95% CI 69.7–85.7), specificity of 93.7% (90.3–96.0), accuracy of 89.9% (86.5–92.5), positive likelihood ratio of 12.5, and a DOR of 55.3. Inter-rater agreement was substantial ( κ  = 0.73; 0.65–0.81), with no systematic discordance (McNemar exact p  = 0.749). Sensitivity was higher for chest than for extremity radiographs (87.3 vs. 63.9%; Fisher's exact p  = 0.010); specificities were comparable. Independent adjudication of the 39 discordant pairs reclassified [X/21] false negatives and [Y/18] false positives as confirmed model errors, [A] as confirmed radiologist omissions or borderline calls, and [B] as legitimately equivocal; the corrected-reference-standard sensitivity and specificity were [S%] and [Sp%] respectively. Conclusion ChatGPT Plus achieved overall diagnostic accuracy close to routine radiologist reports, with substantial agreement and high specificity but lower sensitivity, particularly in extremity radiographs. The model is a plausible supervised educational adjunct or alert application rather than a substitute for expert interpretation; local validation and a human-in-the-loop pathway are prerequisites for any clinical role.

Read PDF

Similar papers

Review Open access Sep 2026

Diagnostic performance of GPT-5.2 ınstant for pediatric elbow fracture detection on radiographs: a prospective single-center diagnostic accuracy study

Multimodal large language models can interpret medical images, but their performance for pediatric elbow radiographs remains uncertain. We evaluated the diagnostic performance of GPT-5.2 Instant as accessed through the ChatGPT web interface during the defined study period. In this prospective, single-cente...

O. Taş, Mehmet Yorgun, R. Aktaş et al. · 0 citations
Open access Sep 2026

Comparison of Multimodal Large Language Models and Oral and Maxillofacial Radiologists in the Detection of Incidental Findings on Panoramic Radiographs: A CBCT-Referenced Diagnostic Accuracy Study

Background/Objectives: This study aimed to compare the diagnostic performance of multimodal (image-capable) large language models (LLMs) and oral and maxillofacial radiologists in detecting nine predefined incidental findings on panoramic radiographs, using a cone-beam computed tomography (CBCT)-derived reference stand...

İsmail Çapar, Utku Cem Hasırcı, Didem Dumanlı Kusay et al. · 0 citations
Open access Aug 2026

Multicenter evaluation of four large language models for automated spine imaging diagnosis

Accurate interpretation of spine imaging is essential for clinical decision-making, yet the diagnostic potential of large language models (LLMs) for radiological report analysis remains inadequately evaluated in terms of sample size, multi-model comparison, reproducibility, and cross-institutional generalisability. Her...

Hao-Lai Liu, Hao Zhang, Hai-Xin Wei et al. · 0 citations
Open access Sep 2026

Laboratory Medicine Decision Support—Beyond Exam Passing: A Blinded 100-Case Text-Based Benchmark of Diagnostic Accuracy, Management Quality, and Safety for ChatGPT, Gemini, and DeepSeek—LLM Decision Support in Laboratory Medicine

Background/Objective: Large language models (LLMs) have shown exam-level performance, yet their reliability and safety in laboratory medicine—where quantitative data interpretation is central—remain insufficiently validated. This study compared the accuracy, interpretive quality, and safety of ChatGPT-5.2, Gemini 3 Pro...

K. Ulutaş, A. Pekmezci · 0 citations
Review Open access Sep 2026

Diagnostic accuracy of point-of-care lung ultrasound compared to chest radiography for identifying community-acquired pneumonia in children: a systematic review and meta-analysis

Chest radiography (CXR) is the conventional imaging standard for diagnosing community-acquired pneumonia (CAP) in children. However, concerns regarding diagnostic variability and ionizing radiation exposure have prompted interest in alternative modalities. Point-of-care ultrasound (POCUS) has emerged as a promising, ra...

A. Shaban, Hany A. Zaki, E. Shaban et al. · 0 citations
Review Open access Sep 2026

How do non-radiologist physicians read and use the radiology report? A French national survey

We developed a national questionnaire to assess radiology reports (RR) physicians’ reading habits and identify their expectations towards RR, focusing on CT and MRI reports. An anonymized 20-item questionnaire was nationally distributed in France via practice and hospital mailing lists and one social media group, and i...

Yasmine Kassab, Matthieu Bailly, Sixtine Brabant et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.