Jul 2026· Journal of Health Sciences and Medicine· Vol 9, pp. 999-1009· 0 citations· 72 references
TL;DR
LLMs/NLPs demonstrated performance comparable to radiologists in many text-based tasks but remain inconsistently evaluated, and GPT-4o had better performance than earlier models.
Abstract
Aims: To systematically review large language model (LLM) and natural language processing (NLP) studies published in first-quartile (Q1) clinical radiology journals, focusing on methodological quality, model implementation, and comparative performance. Methods: A systematic search of PubMed and Scopus was conducted to identify original studies involving LLMs or transformer-based NLP systems published in Q1 clinical radiology journals through June 20, 2025. Eligible studies were screened and assessed for methodological characteristics, including dataset type, involving imaging modality (if any), model used, model accessibility, prompt disclosure, and handling of stochasticity. Human-LLM/NLP and LLM/NLP-LLM/NLP performance comparisons were extracted. Results: Fifty-six studies were included, most published in 2024-2025. Proprietary models such as GPT-4 and GPT-4o were most frequently evaluated. Real-world clinical data were used in 62.5% of studies, but only 10.7% reported a power analysis, and 39.1% addressed stochasticity. Prompt engineering was reported in 41.9% of studies. In 455 human-LLM/NLP comparisons, LLMs/NLPs outperformed humans in 54 cases, while humans outperformed in 79; most results (70.8%) were ties. Among 3,164 valid LLM/NLP-LLM/NLP comparisons, GPT-4o had better performance than earlier models. Conclusion: LLMs/NLPs demonstrated performance comparable to radiologists in many text-based tasks but remain inconsistently evaluated. Methodological limitations, including lack of power analysis, incomplete reporting, and under-addressed stochasticity, hinder robust assessment. Greater transparency, standardized evaluation protocols, and inclusion of diverse clinical settings are essential for reliable integration into radiology practice.
The authors' analysis reveals that LLMs demonstrate promising capabilities in processing textual and visual data related to various liver diseases, including hepatocellular carcinoma, cirrhosis, and non-alcoholic fatty liver disease, but study heterogeneity and significant challenges remain regarding accuracy, reliabil...
T. Suenghataiphorn, Narisara Tribuddharat, Pojsakorn Danpanichkul et al.· Hepatology Forum· 0 citations
Performance of ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers.
Akash Kapoor, Ben Baranker, I. Alter et al.· The Laryngoscope· 0 citations
Enhanced LLMs, particularly DeepSeek-R1, demonstrated robust performance in error detection and correction within real-world Chinese radiology reports, supporting their clinical use for automated quality assurance and integration into workflows to improve reporting accuracy and efficiency.
Jiafeng Zhou, Yuxin Wei, Qian Cai et al.· Journal of Medical Internet...· 0 citations
Background: Systematic reviews summarize research evidence to inform clinical practice guidelines and health policy, but the review process is labour-intensive. Authors often screen thousands of abstracts to identify relevant studies to include in their review. Large Language Models (LLMs) may improve title and abstrac...
Ata Gungor, Adam D. Klotz, K. Solo et al.· McMaster University Medical...· 0 citations
Reporting transparency in radiology and medical imaging LLM studies published in 2025 was inconsistent across reporting items and journals, with substantial deficiencies in some reproducibility-critical elements.
I. Mese, Saime Turgut Gunes, Ozge Coskun et al.· Korean Journal of Radiology· 0 citations
Modern generative large language models (LLMs) are increasingly being evaluated in epilepsy-related clinical tasks, but the evidence remains fragmented and their safe clinical role is uncertain. We conducted a scoping review following the PRISMA-ScR framework, searching PubMed, Embase, and the Web of Science Core Colle...