Skip to content
Open access

A vision transformer deep learning model for assessing pediatric ileocolic intussusception severity using ultrasound images.

Jul 2026 · npj Digital Medicine · 0 citations
Medicine

TL;DR

A Vision Transformer deep learning system is developed and prospectively validated to predict reduction failure from static B-mode ultrasound images to assess pediatric ileocolic intussusception severity, providing an objective, accurate tool to support clinical decision-making and reduce treatment risks.

Abstract

Timely identification of children with ileocolic intussusception likely to fail air-enema reduction is critical to avoid delays and bowel perforation. However, even expert sonographers show inter-observer variability. We developed and prospectively validated a Vision Transformer (ViT) deep learning system to predict reduction failure from static B-mode ultrasound images. This multicenter bidirectional cohort study included 5602 children (4-60 months) who underwent air-enema reduction at 14 Chinese tertiary hospitals (retrospective cohort: 2019-2024). After data augmentation, 10,151 images (8122 training, 2029 validation) were used to train a ViT model for binary classification ("success" vs. "failure"). External validation was performed on a prospective cohort of 190 patients (March-June 2025), with three junior and three senior sonographers independently predicting outcomes. The study was approved by the Ethics Committee of Yijishan Hospital of Wannan Medical University (approval No. 2025-04) and registered with ChiCTR2500098673. The model achieved high internal performance (failure: accuracy 0.880, precision 0.969; success: accuracy 0.970, precision 0.898). In the prospective cohort, the ViT model achieved 93.7% overall accuracy, significantly higher than senior (74.7%) and junior (60.7%) sonographers (p < 0.05). This study innovatively applies ViT to assess pediatric ileocolic intussusception severity, providing an objective, accurate tool to support clinical decision-making and reduce treatment risks.

Read PDF

Similar papers

Open access Aug 2026

Deep learning-based automatic detection of pediatric non-high-density tracheobronchial foreign bodies using chest computed tomography

The developed DL model achieves expert-level accuracy with superior efficiency for detecting pediatric NHDFBs on LDCT, demonstrating strong potential as a rapid, objective decision-support tool to enhance diagnostic workflows, particularly in settings with limited specialist availability.

Junzhong Liu, Qi Wang, Haogang Li et al. · 0 citations
Aug 2026

OA10.1. Deep Learning for Hiatal Hernia Detection on Barium Swallow Images Using a Limited Dataset

Deep learning models can support detection and classification of hiatal hernias using focused AP barium swallow images, even when trained on a limited dataset, and this training approach may enable automated identification and classification while potentially reducing radiation exposure and barium ingestion during diagnostic studies.

B. Borráez-Segura, Ricardo Arango-Slingsby, Lucas Dueñas-Ramirez et al. · 0 citations
Aug 2026

Two-Plane AI-Assisted Renal Ultrasound for Grading Unilateral Hydronephrosis in Infants: A Single-Center Proof-of-Concept Internal Validation Study.

BACKGROUND Renal ultrasound is the first-line imaging method for postnatal hydronephrosis, but its severity grading remains partly reader dependent. Although previous studies have applied artificial intelligence (AI) to pediatric hydronephrosis assessment, the clinical feasibility of a simplified two-still-image workflow in a homogeneous infant cohort remains insufficiently defined. METHODS This retrospective single-center proof-of-concept diagnostic accuracy study was conducted among 186 infants younger than 12 months with unilateral hydronephrosis who had one protocol-compatible transverse and one sagittal renal ultrasound image. The AI workflow classified hydronephrosis as mild, moderate, or severe using a two-branch convolutional neural network (CNN). Model performance was assessed using stratified patient-level five-fold internal validation. The reference standard was expert consensus based on the complete ultrasound examination. Study categories were aligned descriptively with accepted Society for Fetal Urology (SFU) and Urinary Tract Dilation (UTD) severity concepts but were not intended to replace these systems. RESULTS A total of 186 infants were included. Expert consensus classified 79 cases as mild, 62 as moderate, and 45 as severe. The AI-assisted workflow showed exact agreement with expert consensus in 159 of 186 infants, corresponding to 85.5%, with a linear weighted kappa of 0.83. For clinically significant hydronephrosis, defined as moderate or severe disease, the AI workflow showed a sensitivity of 92.5%, a specificity of 86.1%, an accuracy of 89.8%, a positive predictive value of 90.0%, and a negative predictive value of 89.5%. The ordinal-score area under the curve (AUC), which is based on class labels rather than calibrated probability outputs, was 0.92 for clinically significant hydronephrosis and 0.96 for severe hydronephrosis. Because patient-level probability outputs were unavailable, formal calibration, decision curve analysis, threshold adjustment, and individualized risk estimation could not be performed. CONCLUSIONS This single-center proof-of-concept internal validation study suggests that a simplified two-plane AI-assisted workflow could closely match expert consensus grading in selected infants with unilateral hydronephrosis. The workflow should be considered a research-only ordinal class assignment tool. This study does not establish clinical readiness, external generalizability, calibrated risk prediction, or diagnostic superiority over routine reporting. Prospective multicenter validation with preserved probability outputs and formal calibration is needed before any clinical implementation.

Yusuf Atakan Baltrak, Hasan Deliağa · 0 citations
Open access Jul 2026

Development and multicenter external validation of a deep learning model for early screening of thoracic ossification of the ligamentum flavum on routine chest radiographs

Background Thoracic ossification of the ligamentum flavum (TOLF) is frequently underrecognized in its early stage because radiographic abnormalities on routine chest radiographs are often subtle. We aimed to develop and externally validate a deep learning model for opportunistic screening of TOLF using routine chest radiographs. Methods This retrospective multicenter diagnostic study included an internal development cohort from Changzheng Hospital and an independent external validation cohort from South China Hospital. The internal cohort comprised 250 patients with TOLF and 250 control subjects collected between January 2017 and January 2023. The external cohort comprised 150 patients with TOLF and 150 control subjects. TOLF status was established on CT using predefined radiological criteria, whereas frontal and lateral chest radiographs were used only as model inputs. We evaluated multiple backbone architectures, including ResNet101, DenseNet169, Vision Transformer, and Swin Transformer, and additionally explored three dual-view fusion strategies. Model development was performed using 10-fold cross-validation in the internal cohort, and performance was summarized using bootstrap-derived 95% confidence intervals. Human-reader comparison was conducted in the internal cohort. Results In backbone screening within the internal cohort, ResNet101 emerged as the best-performing architecture. After subsequent input-resolution optimization, the final lateral-view ResNet101 model achieved an accuracy of 97.0%, sensitivity of 94.0%, specificity of 100.0%, and an AUC of 0.970. None of the evaluated dual-view fusion strategies outperformed the best single lateral-view model, and the poorer performance of posterior-fusion models was mainly attributable to reduced sensitivity. Compared with experienced spine surgeons and imaging physicians, the internal ResNet101 model showed significantly higher sensitivity and overall accuracy (both p < 0.001). In the external validation cohort, the locked model maintained robust discrimination, with an AUC of 0.954 for frontal radiographs and 0.995 for lateral radiographs. The corresponding accuracy/sensitivity/specificity values were 90.0%/84.7%/95.3% for frontal radiographs and 93.7%/89.3%/98.0% for lateral radiographs. Conclusion A deep learning model based on routine chest radiographs may provide accurate and generalizable screening for TOLF across institutions. The lateral-view model showed the most consistent diagnostic performance, supporting its potential role as an opportunistic screening tool to prompt confirmatory CT evaluation.

Zichuan Wu, Juehan Wang, Xuhong Zhang et al. · 0 citations
Open access Jul 2026

YOLO11-based deep learning system for automated tubal patency classification in hysterosalpingography: a comparative study for clinical decision support.

While the results are promising for a novel application domain, the model's failure on clinically critical minority classes (Bilateral Blockage, Bilateral Patency) means it is not yet suitable for unsupervised clinical use.

Nasreen Jawaid, I. Brohi, Najma Imtiaz Ali et al. · 0 citations
Open access Aug 2026

Intelligent diagnosis of vocal cord lesions via multimodal deep learning: integrating laryngoscopy, voice, and biomarkers

Accurate pre‑operative diagnosis of vocal cord lesions remains challenging. We developed a multimodal deep learning model that integrates laryngoscopic images, voice recordings, and biochemical markers to improve diagnostic accuracy. In this retrospective study (2020–2025), a total of 425 patients were enrolled. Of these, 374 patients treated between January 2020 and May 2025 formed the development cohort (contributing 1,947 aligned multimodal samples), and the remaining 51 consecutive cases treated between June 2025 and December 2025 were reserved as an independent temporal test set. Using an early‑fusion strategy and a Transformer encoder, we built a unified diagnostic framework. Model performance was evaluated by accuracy, average precision, AUC, and loss. During training, the model achieved a sample‑level training accuracy of 98.59% and a validation accuracy of 93.46% (monitored on 198 samples for model selection). For final reporting, all validation and test metrics were computed at the patient level after softmax averaging per patient: the validation set ( n  = 37 patients) yielded an accuracy of 94.59% (35/37) and an AUC of 0.979; the independent temporal test set ( n  = 51 patients) achieved an accuracy of 96.08% (49/51) and an AUC of 0.972, with only two misclassifications. Inference‑masking experiments showed that each modality contributed non‑redundant information; voice masking caused the largest drop (−30.44% accuracy), despite its low static importance weight. Multimodal deep learning may enhance diagnostic accuracy for vocal cord lesions and could help reduce missed diagnoses. Our framework offers a potentially scalable approach for integrating heterogeneous clinical data; however, these findings are preliminary and require confirmation in larger multi‑center cohorts.

Xue Zhao, Shuang Li, Linlin Zheng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.