Skip to content
Review Open access

The use of machine learning models for subdural hematoma detection: a single-arm meta-analysis

Aug 2026 · Neurosurgical review · Vol 49 · 0 citations · 50 references
Medicine

TL;DR

A single arm meta-analysis comparing various DL models for SDH detection depicts potential superiority of U-Net models with respect to sensitivity and precision, based off only 4 pooled U-Net datasets in comparison to the 22 pooled for CNN architectures.

Abstract

Manual evaluation of non-contrast CT scans (NCTS) for detecting subdural hematoma (SDH) is time consuming, potentially inaccurate, and subjective to the expert analyzing them. In recent years, two deep learning (DL) algorithms have been popularly studied in this respect, namely convolutional neural networks (CNN) and U-Net architectures, the latter being a specialized type of CNN. We performed the first meta-analysis comparing various DL models for SDH detection. MEDLINE, Cochrane, Scopus, and Embase databases were searched from inception through December 2025. Studies evaluating ML model performance on an independent test dataset were included. The main outcome measures were sensitivity, specificity, diagnostic odds ratio (DOR), accuracy, and precision of CNN, U-Net, and hybrid DL models. Univariate meta-regression analyses were performed. 30 testing datasets incorporating 67,266 NCTS were included. U-Net demonstrated significantly higher sensitivity (0.916;p = 0.04) and precision (0.983;p = 0.001) while high specificity, DOR, and accuracy values were consistently observed across all DL techniques. Internal testing (p = 0.05) was a borderline significant predictor of high specificity while recent publication year (p < 0.001), U-Net architecture (p = 0.035), and 3D models (p = 0.022) emerged as significant moderators of high precision. The U-Net architecture was also a borderline significant predictor of high DOR (p = 0.049). While this single arm meta-analysis depicts potential superiority of U-Net models with respect to sensitivity and precision, these findings are based off only 4 pooled U-Net datasets in comparison to the 22 pooled for CNN architectures. Future well-powered studies evaluating the U-Net model are necessary to ensure a fair comparison of U-Net architectures to other DL designs before reaching to any definitive conclusions in this respect.

Read PDF

Similar papers

Open access Jul 2026

Deep learning-based detection of acute pancreatitis on abdominal contrast-enhanced CT

DL enabled accurate CECT-based identification of AP in this retrospective multicenter cohort, with performance maintained in an independent external dataset, and showed promising performance for CECT-based acute pancreatitis detection.

Oleksandra Seidel, M. Theis, Sebastian Nowak et al. · 0 citations
Open access Aug 2026

Development of a deep learning model for detecting and measuring gallstones on computed tomography images.

The developed model achieves high sensitivity and precise automated gallstone segmentation on CT images and achieves overall sensitivities of 97.2%, 97.5%, 95.2%, 89.2%, and 98.2% across the training, validation, internal test, hold-out, and AMOS datasets.

Yue Gao, Yaofeng Zhang, Xiao-Dong Zhang et al. · 0 citations
Review Open access Aug 2026

Deep learning model for automatic detection of incidental adrenal abnormalities on low-dose computed tomography images

To investigate the feasibility of employing deep learning models for automated segmentation and classification of adrenal incidental abnormalities on low-dose CT images. Four distinct CT cohorts were retrospectively collected for deep learning models development (cohort A, n  = 2574; cohort B, n  = 1205), internal evaluation (cohort C, n  = 3681), and external evaluation (cohort D, n  = 779). Two experienced uroradiologists independently reviewed the CT images and labeled the adrenal glands as normal or abnormal based on predefined criteria encompassing both density and morphological abnormalities, with any discrepancies resolved through consultation. The model development cohorts were divided into a training set, a validation set, and a test set. Deep learning models for segmentation and classification were trained and evaluated on internal and external sets, with the dice similarity coefficient (DSC), area under precision–recall curves (AUPRC), and area under receiver operating characteristic curves (AUROC) as evaluation metrics. Adrenal descriptions from radiology reports were extracted to compare with the model’s performance. For adrenal gland segmentation, the DSC values for the test set, internal validation cohort, and external validation cohort were 0.839 (IQR: 0.783–0.871), 0.870 (IQR: 0.819–0.902), and 0.799 (IQR: 0.729–0.849), respectively. For adrenal gland classification, the AI model achieved AUPRC values of 0.913, 0.753, and 0.927 in the test set, internal validation cohort, and external validation cohort, respectively, outperforming routine radiology reporting (AUPRC: 0.809, 0.708, 0.591; all P  < 0.05). Corresponding AUROC values were 0.956, 0.942, and 0.977 for the AI model, which also outperformed routine radiology reporting (AUROC: 0.889, 0.705, 0.551; all P  < 0.05). The deep learning models showed promise in automated adrenal segmentation and classification, highlighting AI’s potential to improve detection of adrenal abnormalities in LDCT scans. This study has been registered on ClinicalTrials.gov on August 25, 2025, with the unique identifier NCT07198152.

Kexin Wang, He Wang, Shiwei Chen et al. · 0 citations
Open access Jul 2026

A deep learning model for the interpretable identification of pulmonary thromboembolism from computed tomography pulmonary angiography

Objective The rapid identification of pulmonary thromboembolism (PTE) on computed tomography pulmonary angiography (CTPA) is vital but labor-intensive, often leading to diagnostic delays. We aimed to construct and evaluate a YOLOv11 object detection algorithm capable of automatically highlighting intraluminal filling defects to expedite emergency radiological workflows. Methods A retrospective analysis was conducted on CTPA scans from multiple centers. The dataset was divided into a primary internal cohort (n = 1,368) for model derivation and testing, alongside an independent external cohort (n = 98) to assess generalizability. The diagnostic efficacy of the YOLOv11 architecture was quantified using the area under the receiver operating characteristic curve (AUC), sensitivity, and specificity. Additionally, gradient-weighted class activation mapping (Grad-CAM) was applied to map the spatial distribution of the model's focus, ensuring clinical transparency. Results During internal testing, the proposed framework yielded an AUC of 0.777 [95% confidence interval (CI): 0.765–0.788], corresponding to a sensitivity of 74.53% and a specificity of 64.26%. When applied to the external cohort, the algorithm's discriminative ability remained consistent with an AUC of 0.778 (95% CI: 0.749–0.806). Notably, the external sensitivity reached 86.75% (specificity: 54.46%). Visual assessments via Grad-CAM saliency maps confirmed that the model accurately localized embolic occlusions within the complex pulmonary arterial tree. Conclusion Utilizing the YOLOv11 architecture for automated CTPA analysis yields a highly sensitive and visually interpretable screening mechanism. This artificial intelligence-assisted approach holds substantial promise for reducing missed diagnoses and accelerating patient triage in acute clinical settings.

Yemei Li, Sunyu Chen, Guangkun Chen et al. · 0 citations
Review Aug 2026

Deep learning versus radiologists for acute aortic dissection on CT: A systematic review and network meta-analysis.

Acute aortic dissection (AAD) is a time-sensitive cardiovascular emergency in which delayed recognition remains associated with poor early outcomes, and deep learning (DL) applied to computed tomography (CT) is increasingly proposed for triage support. We synthesized DL detection accuracy on CT, compared DL with radiologists head-to-head, and secondarily pooled CT segmentation accuracy. This Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA)-compliant review searched PubMed, Embase, and Web of Science through February 18, 2026. Detection studies were pooled with bivariate random-effects models; head-to-head comparisons were synthesized with contrast-based network meta-analysis; segmentation Sørensen-Dice coefficients were analyzed with three-level random-effects models. Twenty-three studies were included. For non-contrast CT detection (10 studies; 18 cohorts; 54 algorithm-cohort datasets), pooled sensitivity was 0.91 (95 % confidence interval [CI], 0.87-0.94) and specificity 0.90 (95 % CI, 0.85-0.94). In direct comparisons (5 studies; 11 comparisons), radiologist-to-DL relative sensitivity was 0.72 (95 % CI, 0.57-0.91) and relative specificity 0.88 (95 % CI, 0.83-0.93). For the secondary segmentation synthesis (16 studies; 88 datasets), pooled Dice was 0.863 (95 % CI, 0.795-0.931), lower for true- and false-lumen than for whole-aorta delineation. DL showed encouraging detection accuracy and higher sensitivity and specificity than the radiologist comparators in five head-to-head studies read under experimental conditions, supporting assistive triage rather than replacement; lumen-level segmentation remains more challenging than whole-aorta contouring. Certainty was low - datasets were predominantly case-control with a median prevalence of 50 %, heterogeneity was substantial and small-study effects pronounced - so these estimates are experimental upper bounds requiring prospective confirmation.

Ting-Wei Wang, Jia-Sheng Hong, Ho-Ren Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.