Skip to content
Open access

Pediatric vs. Adult Pneumonia Detection: Quantifying Age-related Generalization Gaps in Zero-shot Multimodal Large Language Models.

Aug 2026 · Academic Radiology · 0 citations · 15 references
Medicine

TL;DR

Zero-shot multimodal LLMs show large age-related generalization gaps and clinically relevant error asymmetries in pediatric chest radiography, whereas a domain-trained convolutional neural network remains robust within its training domain.

Abstract

Rationale

AND

Objectives

Multimodal large language models (LLMs) are increasingly applied to image-based radiology tasks, but their diagnostic accuracy across clinically distinct populations remains poorly characterized. We quantified age-related differences in zero-shot LLM performance for pneumonia detection on pediatric vs. adult chest radiographs and compared generalization gaps with a domain-trained convolutional neural network (CNN) baseline.

Materials And Methods

GPT-5.2 (OpenAI), Claude Opus 4.5 (Anthropic), and Gemini 2.5 Pro (Google) were evaluated zero-shot on balanced pediatric and adult test sets of frontal chest radiographs (n = 1000 each; 500 pneumonia/500 normal). Cohort-specific InceptionV3 CNNs were trained on the remaining development-pool images (pediatric n = 4715; adult n = 13,863) and evaluated on the same test sets. Performance was assessed using the Matthews correlation coefficient (MCC) with 95% bootstrap confidence intervals (CIs); domain shift was quantified as Δ = Adult - Pediatric.

Results

In pediatrics, the CNN outperformed all LLMs (MCC 0.799, 95% CI 0.766-0.832) vs. GPT-5.2 (0.484, 0.436-0.532), Claude Opus 4.5 (0.470, 0.418-0.521), and Gemini 2.5 Pro (0.272, 0.224-0.316). In adults, all models improved, but the CNN remained best (MCC 0.850, 0.816-0.882). Age-related gains were larger for LLMs (ΔMCC +0.220 [95% CI 0.160-0.281] to +0.466 [0.407-0.525]) than for the CNN (ΔMCC +0.051 [0.005-0.098]), driven mainly by specificity increases.

Conclusion

Zero-shot multimodal LLMs show large age-related generalization gaps and clinically relevant error asymmetries in pediatric chest radiography, whereas a domain-trained CNN remains robust within its training domain. Rigorous subgroup evaluation, including pediatric populations, is essential before clinical deployment of multimodal LLMs.

Read PDF

Similar papers

Open access Aug 2026

Multicenter cross domain evaluation of CNNs and vision transformers trained on adult data for pediatric pneumonia screening

Large-scale adult chest radiograph datasets are commonly used to develop deep learning models for pneumonia detection, whereas pediatric pneumonia datasets remain smaller, more heterogeneous, and less widely available. Because pediatric chest radiographs differ from adult radiographs in anatomical proportions, diseas...

Yutao Li, Junghun Kim, Sang-Il Choi · 0 citations
Open access Jul 2026

Integrating Local and Global Representation Learning for Pediatric Pneumonia Detection: A Hybrid CNN–Transformer Ensemble Framework

Background/Objectives: Pneumonia remains a leading cause of childhood morbidity and mortality worldwide. Accurate interpretation of pediatric chest radiographs is challenging because of anatomical variability, subtle radiographic findings, and inter-observer variability. This study evaluates different CNN–Transformer e...

Ece Meltem Yalçın, Hayriye Tanyıldız, Serpil Aslan et al. · 1 citation
Open access Sep 2026

Automated chest X-ray disease screening using large language models and deep convolutional neural networks on the MIMIC-CXR dataset

It is demonstrated that LLMs can be effectively employed to generate supervision labels for medical imaging tasks and that the proposed approach offers a scalable and low-cost solution for preliminary disease screening, particularly in healthcare environments with limited expert availability.

Qing-Yuan Zhang, Pardeep Vasudev, Kezhi Li et al. · 0 citations
Open access Sep 2026

COMPARISON OF PERFORMANCE AND COMPUTATIONAL COMPLEXITY OF CNN AND RESNET50 FOR PNEUMONIA CLASSIFICATION

Pneumonia is an acute infection of the lung tissue and the leading cause of death among children under five years of age worldwide. Diagnosis based on chest X-ray images is considered prone to misinterpretation, particularly in mild cases that appear similar to normal lung conditions. CNN is one of the most widely used...

Tam Pran Noto Noto, Supatman Supatman · 0 citations
Conference Open access 2025

Intelligent Pneumonia Screening: A Convolutional Neural Network Approach Using Chest Radiographs

: Pneumonia continues to be a major source of morbidity and mortality globally, especially in developing countries where a shortage of specialists makes radiological assessment challenging. Patients' survival and appropriate treatment depend on a timely and accurate diagnosis. This study examines a deep learning techni...

M. Devi, Tanya, Aradhya Mittal et al. · 0 citations
Open access Sep 2026

Improving Diagnostic Sensitivity in Imbalanced Oral Cancer Image Classification: A Comparative Study of CNN and Transformer Architectures

Simple Summary Early detection of oral cancer is essential for improving patient survival, yet artificial intelligence models often struggle because available clinical image datasets are small and highly imbalanced, with relatively few malignant cases. This study investigates whether class imbalance can be mitigated th...

Pablo Ormeño-Arriagada, Valentina Zúñiga, Carlos A. Toro et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.