Skip to content
Open access

Deep learning-based automatic detection of pediatric non-high-density tracheobronchial foreign bodies using chest computed tomography

Aug 2026 · Scientific Reports · 0 citations

TL;DR

The developed DL model achieves expert-level accuracy with superior efficiency for detecting pediatric NHDFBs on LDCT, demonstrating strong potential as a rapid, objective decision-support tool to enhance diagnostic workflows, particularly in settings with limited specialist availability.

Abstract

Foreign body aspiration (FBA) of non-high-density objects (NHDFBs) in children is a critical pediatric emergency, posing risks of airway obstruction and requiring prompt diagnosis, which currently relies on clinician experience with low-dose computed tomography (LDCT). This study aimed to develop and validate a deep learning (DL) model for the automated detection of tracheobronchial NHDFBs in pediatric LDCT scans. A retrospective, multicenter cohort of 600 children with suspected FBA was utilized, with bronchoscopic confirmation as the gold standard. A ResUnet-based model was trained and evaluated on internal and external validation sets, with its performance systematically compared against junior and senior radiologists. The DL model achieved foreign body detection rates comparable to senior radiologists across training, internal test and external validation cohorts without significant intergroup differences, whereas both outperformed junior radiologists significantly (all p  < 0.05). The DL model exhibited drastically shorter reading time (17.5 ± 2.4 s) than senior radiologists (80.2 ± 13.1 s) and junior radiologists (109.3 ± 16.8 s, all p  < 0.05). The model maintained stable high diagnostic performance in all cohorts. Its sensitivity was significantly higher than junior radiologists in both validation sets (all p  < 0.05), while sensitivity, specificity, PPV and NPV showed no statistical disparities between the DL model and senior radiologists. The DL model yielded AUC values of 0.93, 0.89 and 0.85 in the three cohorts, which were statistically equivalent to those of senior radiologists (0.94, 0.95, 0.90), and substantially superior to junior radiologists (0.82, 0.83, 0.79). The developed DL model achieves expert-level accuracy with superior efficiency for detecting pediatric NHDFBs on LDCT, demonstrating strong potential as a rapid, objective decision-support tool to enhance diagnostic workflows, particularly in settings with limited specialist availability.

Read PDF

Similar papers

Open access Jul 2026

Deep learning-based detection of acute pancreatitis on abdominal contrast-enhanced CT

DL enabled accurate CECT-based identification of AP in this retrospective multicenter cohort, with performance maintained in an independent external dataset, and showed promising performance for CECT-based acute pancreatitis detection.

Oleksandra Seidel, M. Theis, Sebastian Nowak et al. · 0 citations
Open access Jul 2026

A CT-based deep learning model for the automated risk stratification of refractory Mycoplasma pneumoniae pneumonia in children.

BACKGROUND The accurate identification of children with refractory Mycoplasma pneumoniae pneumonia (RMPP) remains challenging. This study aimed to develop a transformer-based model utilizing clinically indicated chest computed tomography (CT) to stratify pediatric RMPP risk at a critical decision point. METHODS Non-contrast chest CT data from a multicenter retrospective cohort of 1224 pediatric patients with Mycoplasma pneumoniae pneumonia who underwent clinically indicated CT were used to develop a transformer-based deep learning framework (trans-DLF). The primary cohort comprised training (n = 506), validation (n = 140), and internal testing (n = 139) cohorts, with two independent external cohorts (n = 331 and n = 108) used to evaluate generalizability. Model performance was assessed by the area under the receiver operating characteristic curve (AUC) and compared against a three-dimensional convolutional neural network (3D-CNN), a clinical model, and a multimodal nomogram. Interpretability was examined using gradient-weighted class activation mapping (Grad-CAM). RESULTS The median age was 6.83 years (interquartile range, 5.0-8.6 years), and 609 (49.8%) were male. The trans-DLF demonstrated strong performance across all cohorts: training (AUC 0.97; 95% confidence interval [CI], 0.96-0.98), validation (0.91; 0.86-0.96), internal testing (0.90; 0.85-0.95), and external testing (0.89; 0.84-0.94 and 0.89; 0.82-0.95). It significantly outperformed the clinical model (p < 0.001), while its AUCs were not significantly different from those of the multimodal nomogram. The model maintained good performance in outpatient settings (AUC 0.87) with good calibration and net clinical benefit. Grad-CAM suggested that predictions were influenced by clinically meaningful features, particularly consolidations. CONCLUSION The trans-DLF provides a streamlined and efficient approach to RMPP risk assessment in children who have already undergone clinically indicated chest CT and may support timely, evidence-based decision-making without additional tests.

Zhoumeng Ying, Ge Hu, Jing Li et al. · 0 citations
Open access Sep 2026

Deep learning on contrast-enhanced computed tomography for parotid tumor classification: Providing crucial diagnostic information beyond clinical and radiologic evaluation.

PURPOSE Preoperative differentiation between benign and malignant parotid tumors (PTs) remains challenging despite clinical examination, cross-sectional imaging, and biopsy. Accurate malignancy assessment is crucial to optimize surgical planning and minimize morbidity. METHODS We retrospectively analyzed 66 patients who underwent parotidectomy between 2008 and 2024, including 33 with malignant and 33 with benign PTs matched for demographics. All patients had contrast-enhanced computed tomography (CE-CT), and segmented tumor volumes were evaluated using a three-dimensional convolutional neural network. Diagnostic performance was compared with standard clinical and radiologic workup. RESULTS Standard clinical and radiologic workup achieved a sensitivity of 60.6 %, increasing to 69.7 % with fine needle aspiration cytology (FNAC) or core needle biopsy (CNB). The deep learning model achieved an area under the ROC curve of 0.94, with both sensitivity and specificity exceeding 90 % at optimized thresholds, outperforming conventional diagnostics. CONCLUSION Deep learning applied to CE-CT demonstrated strong diagnostic performance for the preoperative classification of PTs in this cohort and may serve as a powerful non-invasive adjunct to standard diagnostic modalities without adding procedural burden to the diagnostic workup.

Unknown authors · 0 citations
Open access Aug 2026

Deep Learning-Assisted MRI for Differentiating Parotid Gland Tumors: Comparison with FNAB and Radiologic Assessment

Background/Objectives: To evaluate the diagnostic performance of an MRI-based deep learning (DL) model for differentiating benign and malignant parotid gland tumors and to compare its performance with radiologic assessment and fine-needle aspiration biopsy (FNAB), using histopathology as the reference standard. Methods: This retrospective single-center study included 144 consecutive patients with histopathologically confirmed parotid gland tumors who underwent parotidectomy between January 2020 and December 2024. Preoperative MRI examinations were analyzed using a ResNet50-based convolutional neural network incorporating a Convolutional Block Attention Module (CBAM). Model performance was evaluated using patient-level stratified 5-fold cross-validation and compared with MRI and FNAB using sensitivity, specificity, accuracy, positive predictive value (PPV), and negative predictive value (NPV). Results: The DL model achieved a sensitivity of 79.3% (95% CI, 61.6–90.2), specificity of 93.9% (95% CI, 87.9–96.9), and an overall accuracy of 91.0% (95% CI, 85.3–94.5). MRI demonstrated a sensitivity of 86.2% and specificity of 93.0%, whereas FNAB achieved a sensitivity of 82.8% and specificity of 96.5%. The DL model showed consistent performance across the validation folds. Conclusions: MRI-based deep learning demonstrated high diagnostic performance for the preoperative classification of parotid gland tumors and may serve as a complementary decision-support tool alongside MRI and FNAB. Although promising, these findings are limited by the retrospective single-center design and the lack of external validation. Prospective multicenter studies are required before routine clinical implementation.

Servet Erdemes, Ö. Türk, Mahmut Ağırtmış et al. · 0 citations
Review Open access Aug 2026

Deep learning model for automatic detection of incidental adrenal abnormalities on low-dose computed tomography images

To investigate the feasibility of employing deep learning models for automated segmentation and classification of adrenal incidental abnormalities on low-dose CT images. Four distinct CT cohorts were retrospectively collected for deep learning models development (cohort A, n  = 2574; cohort B, n  = 1205), internal evaluation (cohort C, n  = 3681), and external evaluation (cohort D, n  = 779). Two experienced uroradiologists independently reviewed the CT images and labeled the adrenal glands as normal or abnormal based on predefined criteria encompassing both density and morphological abnormalities, with any discrepancies resolved through consultation. The model development cohorts were divided into a training set, a validation set, and a test set. Deep learning models for segmentation and classification were trained and evaluated on internal and external sets, with the dice similarity coefficient (DSC), area under precision–recall curves (AUPRC), and area under receiver operating characteristic curves (AUROC) as evaluation metrics. Adrenal descriptions from radiology reports were extracted to compare with the model’s performance. For adrenal gland segmentation, the DSC values for the test set, internal validation cohort, and external validation cohort were 0.839 (IQR: 0.783–0.871), 0.870 (IQR: 0.819–0.902), and 0.799 (IQR: 0.729–0.849), respectively. For adrenal gland classification, the AI model achieved AUPRC values of 0.913, 0.753, and 0.927 in the test set, internal validation cohort, and external validation cohort, respectively, outperforming routine radiology reporting (AUPRC: 0.809, 0.708, 0.591; all P  < 0.05). Corresponding AUROC values were 0.956, 0.942, and 0.977 for the AI model, which also outperformed routine radiology reporting (AUROC: 0.889, 0.705, 0.551; all P  < 0.05). The deep learning models showed promise in automated adrenal segmentation and classification, highlighting AI’s potential to improve detection of adrenal abnormalities in LDCT scans. This study has been registered on ClinicalTrials.gov on August 25, 2025, with the unique identifier NCT07198152.

Kexin Wang, He Wang, Shiwei Chen et al. · 0 citations
Open access Jul 2026

Development and multicenter external validation of a deep learning model for early screening of thoracic ossification of the ligamentum flavum on routine chest radiographs

Background Thoracic ossification of the ligamentum flavum (TOLF) is frequently underrecognized in its early stage because radiographic abnormalities on routine chest radiographs are often subtle. We aimed to develop and externally validate a deep learning model for opportunistic screening of TOLF using routine chest radiographs. Methods This retrospective multicenter diagnostic study included an internal development cohort from Changzheng Hospital and an independent external validation cohort from South China Hospital. The internal cohort comprised 250 patients with TOLF and 250 control subjects collected between January 2017 and January 2023. The external cohort comprised 150 patients with TOLF and 150 control subjects. TOLF status was established on CT using predefined radiological criteria, whereas frontal and lateral chest radiographs were used only as model inputs. We evaluated multiple backbone architectures, including ResNet101, DenseNet169, Vision Transformer, and Swin Transformer, and additionally explored three dual-view fusion strategies. Model development was performed using 10-fold cross-validation in the internal cohort, and performance was summarized using bootstrap-derived 95% confidence intervals. Human-reader comparison was conducted in the internal cohort. Results In backbone screening within the internal cohort, ResNet101 emerged as the best-performing architecture. After subsequent input-resolution optimization, the final lateral-view ResNet101 model achieved an accuracy of 97.0%, sensitivity of 94.0%, specificity of 100.0%, and an AUC of 0.970. None of the evaluated dual-view fusion strategies outperformed the best single lateral-view model, and the poorer performance of posterior-fusion models was mainly attributable to reduced sensitivity. Compared with experienced spine surgeons and imaging physicians, the internal ResNet101 model showed significantly higher sensitivity and overall accuracy (both p < 0.001). In the external validation cohort, the locked model maintained robust discrimination, with an AUC of 0.954 for frontal radiographs and 0.995 for lateral radiographs. The corresponding accuracy/sensitivity/specificity values were 90.0%/84.7%/95.3% for frontal radiographs and 93.7%/89.3%/98.0% for lateral radiographs. Conclusion A deep learning model based on routine chest radiographs may provide accurate and generalizable screening for TOLF across institutions. The lateral-view model showed the most consistent diagnostic performance, supporting its potential role as an opportunistic screening tool to prompt confirmatory CT evaluation.

Zichuan Wu, Juehan Wang, Xuhong Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.