TPTh 2.13 Machine Learning for the Preoperative Prediction of Perforated Appendicitis: A Systematic Review and Diagnostic Test Accuracy Meta-Analysis
Perforated appendicitis increases morbidity and is difficult to predict preoperatively. This systematic review and diagnostic meta-analysis evaluated machine learning (ML) models for their preoperative prediction. On November 19, 2025, PubMed, Scopus, Web of Science, Embase, Google Scholar, and references were searched per PRISMA 2020. Risk of bias was assessed with PROBAST. Two reviewers screened and extracted data. A bivariate random-effects meta-analysis estimated pooled sensitivity, specificity, diagnostic odds ratio (DOR), and HSROC; remaining studies were synthesized descriptively, with sensitivity analyses for robustness. Five retrospective studies (n=6,224) with mean ages 10.7–43 years. The pooled sensitivity across the three eligible studies for meta-analysis was 0.86 (95% CI: 0.66–0.95) and the pooled specificity was 0.96 (95% CI: 0.33–0.99). The DOR was markedly elevated (149.7; 95% CI: 4.9–4528.3), indicating strong overall discriminatory ability of ML approaches. A negative correlation was observed between sensitivity and specificity (random-effects correlation = –0.75), consistent with threshold effects across studies. HSROC model parameters (θ = 0.75, λ = 5.04, β = 1.18) confirmed high overall accuracy and illustrated variability in diagnostic thresholds among included models. Sensitivity analyses yielded consistent results, supporting the stability of the findings. Risk of bias unclear (retrospective design, limited validation); applicability high in 60%, unclear in 40%. ML shows promise for preoperative detection of perforated appendicitis, but variability and limited validation require prospective studies before clinical use. Strengthening external validation and standardizing methodologies will be essential for safe clinical integration.