Multicenter cross domain evaluation of CNNs and vision transformers trained on adult data for pediatric pneumonia screening
Abstract
Large-scale adult chest radiograph datasets are commonly used to develop deep learning models for pneumonia detection, whereas pediatric pneumonia datasets remain smaller, more heterogeneous, and less widely available. Because pediatric chest radiographs differ from adult radiographs in anatomical proportions, disease presentation, and acquisition characteristics, models trained on adult data may not generalize reliably to pediatric populations. This study therefore evaluated the cross-domain performance, stability, and interpretability of convolutional neural networks (CNNs) and vision transformers (ViTs) for pediatric pneumonia screening. A frontal-only adult CheXpert AP/PA cohort was used for training, and three independent pediatric datasets—Kaggle Pediatric, NIH Pediatric Expanded, and VinDr-PCXR—were used for external testing. Three architectures—EfficientNet-B0, ConvNeXt-Tiny, and ViT-Base-16—were evaluated across five independent runs, with interpretability assessed using Score-CAM and quantitative attention-overlap analysis against radiologist-annotated abnormal regions on VinDr-PCXR. ConvNeXt-Tiny achieved the highest mean F1-score on Kaggle Pediatric and NIH Pediatric Expanded. On the more challenging VinDr-PCXR dataset, ViT-Base-16 showed markedly lower sensitivity than ConvNeXt-Tiny (18.9 ± 9.7% vs. 53.3 ± 24.6%) and weaker score-CAM overlap with annotated abnormal regions. Among the evaluated architectures, ConvNeXt-Tiny showed comparatively favorable and stable cross-domain behavior, although no architecture eliminated the adult-to-pediatric domain shift. These findings underscore the need for further pediatric-specific external validation before clinical translation.