Skip to content
Review Open access

Large language models and natural language processing applications in radiology: a systematic review of Q1 journal studies

Jul 2026 · Journal of Health Sciences and Medicine · Vol 9, pp. 999-1009 · 0 citations · 72 references

TL;DR

LLMs/NLPs demonstrated performance comparable to radiologists in many text-based tasks but remain inconsistently evaluated, and GPT-4o had better performance than earlier models.

Abstract

Aims: To systematically review large language model (LLM) and natural language processing (NLP) studies published in first-quartile (Q1) clinical radiology journals, focusing on methodological quality, model implementation, and comparative performance. Methods: A systematic search of PubMed and Scopus was conducted to identify original studies involving LLMs or transformer-based NLP systems published in Q1 clinical radiology journals through June 20, 2025. Eligible studies were screened and assessed for methodological characteristics, including dataset type, involving imaging modality (if any), model used, model accessibility, prompt disclosure, and handling of stochasticity. Human-LLM/NLP and LLM/NLP-LLM/NLP performance comparisons were extracted. Results: Fifty-six studies were included, most published in 2024-2025. Proprietary models such as GPT-4 and GPT-4o were most frequently evaluated. Real-world clinical data were used in 62.5% of studies, but only 10.7% reported a power analysis, and 39.1% addressed stochasticity. Prompt engineering was reported in 41.9% of studies. In 455 human-LLM/NLP comparisons, LLMs/NLPs outperformed humans in 54 cases, while humans outperformed in 79; most results (70.8%) were ties. Among 3,164 valid LLM/NLP-LLM/NLP comparisons, GPT-4o had better performance than earlier models. Conclusion: LLMs/NLPs demonstrated performance comparable to radiologists in many text-based tasks but remain inconsistently evaluated. Methodological limitations, including lack of power analysis, incomplete reporting, and under-addressed stochasticity, hinder robust assessment. Greater transparency, standardized evaluation protocols, and inclusion of diverse clinical settings are essential for reliable integration into radiology practice.

Read PDF

Similar papers

Review Open access Apr 2026

Large language models in hepatology: A systematic review

The authors' analysis reveals that LLMs demonstrate promising capabilities in processing textual and visual data related to various liver diseases, including hepatocellular carcinoma, cirrhosis, and non-alcoholic fatty liver disease, but study heterogeneity and significant challenges remain regarding accuracy, reliabil...

T. Suenghataiphorn, Narisara Tribuddharat, Pojsakorn Danpanichkul et al. · 0 citations
Review Aug 2026

ENTGPT: Applying Large Language Models to Systematic Review Screening With the Novel STARR Protocol.

Performance of ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers.

Akash Kapoor, Ben Baranker, I. Alter et al. · 0 citations
Review Open access Aug 2026

Error Detection and Correction in Chinese Radiology Reports Using Large Language Models: Real-World Clinical Validation Study

Enhanced LLMs, particularly DeepSeek-R1, demonstrated robust performance in error detection and correction within real-world Chinese radiology reports, supporting their clinical use for automated quality assurance and integration into workflows to improve reporting accuracy and efficiency.

Jiafeng Zhou, Yuxin Wei, Qian Cai et al. · 0 citations
Review Open access Jul 2026

Performance of large language models on screening titles and abstracts of urology-related systematic reviews

Background: Systematic reviews summarize research evidence to inform clinical practice guidelines and health policy, but the review process is labour-intensive. Authors often screen thousands of abstracts to identify relevant studies to include in their review. Large Language Models (LLMs) may improve title and abstrac...

Ata Gungor, Adam D. Klotz, K. Solo et al. · 0 citations
Review Open access Aug 2026

Reporting Quality of Large Language Model Studies: A Cross-Sectional Audit of High-Ranking Radiology and Medical Imaging Journals

Reporting transparency in radiology and medical imaging LLM studies published in 2025 was inconsistent across reporting items and journals, with substantial deficiencies in some reproducibility-critical elements.

I. Mese, Saime Turgut Gunes, Ozge Coskun et al. · 0 citations
Review Open access Sep 2026

Modern generative large language models in epilepsy care: a scoping review of current applications, challenges, and future directions

Modern generative large language models (LLMs) are increasingly being evaluated in epilepsy-related clinical tasks, but the evidence remains fragmented and their safe clinical role is uncertain. We conducted a scoping review following the PRISMA-ScR framework, searching PubMed, Embase, and the Web of Science Core Colle...

Shi-Hao Ge, Yue-Qian Sun, Qun Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.