Aug 2026· Radiology· Vol 320 2, pp.
e253238
· 1 citation· 15 references
Medicine
TL;DR
LLMs generated radiology-relevant indications from clinical notes that were more comprehensive and factual than clinician indications, and when generated by the proprietary LLM, were ranked most useful in protocoling and imaging interpretation.
The addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, but this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.
Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al.· Journal of Medical Internet...· 0 citations
BACKGROUND
Chest radiography remains the most widely used imaging modality worldwide; however, its interpretation is inherently challenging because of overlapping anatomical structures and subtle findings. Recent advances in multimodal large language models (LLMs) have enabled automated radiology report generation, yet their clinical performance relative to domain-specific medical AI systems remains insufficiently validated.
OBJECTIVES
This study aimed to evaluate the performance and clinical applicability of a domain-specific multimodal AI model (M4CXR) compared with a general-purpose LLM (ChatGPT-4o) for chest radiograph interpretation.
METHODS
In this retrospective study, 500 anonymized chest radiographs from a single tertiary care center were analyzed. Four board-certified radiologists independently evaluated AI-generated reports from both models. Key outcomes included key finding detection (categorized as complete, partial, or inconsistent), report generation time, and report discrepancies assessed using the RADPEER scoring system. Agreement between original and M4CXR-assisted RADPEER scores was assessed using intraclass correlation coefficients and weighted Cohen's kappa. Statistical analyses included paired t-tests, and chi-square tests.
RESULTS
M4CXR demonstrated significantly higher report consistency than GPT-4o, with complete concordance observed in 55.8% versus 19.8% of cases, and lower inconsistency rates (25.2% vs. 46.4%, P<.001). The use of M4CXR significantly reduced report generation time compared with unaided interpretation (16.3 ± 12.9 s vs. 179.2 ± 50.4 s, P<.001). RADPEER-based discrepancy analysis revealed no significant differences between original and AI-assisted interpretations. Agreement between original and M4CXR-assisted RADPEER scores showed good reliability (ICC = 0.701), and weighted kappa analysis showed substantial agreement (κw = 0.652).
CONCLUSIONS
Domain-specific multimodal AI model evaluated in this study demonstrated higher diagnostic consistency than the general-purpose LLM evaluated under the study conditions. These findings suggest the potential of specialized AI models as viable assistive tools, while highlighting the complementary utility of general-purpose LLMs in broader clinical contexts. Future integration should prioritize human-AI collaboration and prospective multi-center validation.
Tae-Hoon Kim, J. Hong, Jihun Hyun et al.· BMC Medical Imaging· 0 citations
BACKGROUND
Large language models (LLMs) show promise for converting complex radiology reports into patient-centric language, but inherent output instability may limit clinical application.
OBJECTIVES
To quantitatively assess the translational accuracy, error rates, and instability of various LLMs when generating patient-centric radiology reports, and evaluate demographic influences on report readability.
MATERIALS AND METHODS
This retrospective study evaluated 320 de-identified radiology reports processed by three LLMs using a two-stage (baseline and optimized) prompt engineering strategy. Two senior radiologists evaluated medical accuracy, completeness, and recommendation suitability. Readability was evaluated by 16 non-medical participants stratified by age and education.
RESULTS
Professional radiological evaluation revealed that all tested models exhibited inherent instability, omitted information, and tended to generate risk-averse, generalized clinical recommendations. To address these limitations, optimized structured prompts significantly reduced model output variance and improved translational accuracy, with particularly prominent effects observed in DeepSeek-R1 and ChatGPT-4.0. Overall, large language models significantly enhanced the readability of radiology reports (P < 0.05), with DeepSeek-R1 achieving the best performance. However, patients' self-reported comprehension of the reports was affected by demographic characteristics.
CONCLUSION
Large language models can effectively improve the readability of radiology reports, yet all such models inherently suffer from output instability and information omission. Optimized structured prompting can substantially reduce the variability of model outputs and improve the accuracy of medical text translation. Nevertheless, LLMs should currently be strictly confined to human-supervised auxiliary tools rather than applied as standalone clinical solutions.
Yun Mao, Chunyan Wang, Wei Wang et al.· Academic Radiology· 0 citations
A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.
Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al.· Indian Journal of Computer S...· 0 citations
Although users rated RAG-enabled GPT-4o agent-generated documentation summaries favorably on several quality domains, they frequently lacked thoroughness and occasionally contained treatment-relevant errors.
Ko Un Park, Bergen K. Sather, A. Shah et al.· Annals of Surgical Oncology· 0 citations
This work introduces MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm.
Lai Wei, Yu-Chao Chen, Zhenbiao Cao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.