Machine learning-based risk prediction model for postoperative abdominal distension after spinal surgery under general anesthesia
Abstract
Early identification of patients at risk of postoperative abdominal distension after spinal surgery remains clinically difficult. Machine learning may improve risk prediction by integrating routinely available perioperative information. This study aimed to develop and internally validate an interpretable machine learning-based predictive model for postoperative abdominal distension after spinal surgery under general anesthesia and to identify clinically relevant predictors for early nursing risk stratification. This retrospective single-center prediction model study included 1,234 patients who underwent spinal surgery under general anesthesia. Seventeen routinely recorded perioperative and early postoperative variables were evaluated, including postoperative ADL category(ADL), PCA, and postoperative antibiotic use (yes/no). The data were split using stratified sampling into training (80%) and internal test (20%) sets. Missing-value imputation, scaling, five-fold cross-validated grid search, and model fitting were performed within the training workflow. Logistic regression(LR), RF, support vector machine(SVM), and gradient boosting (GB)models were evaluated using discrimination, calibration, Brier score, decision curve analysis, and permutation importance. Exact TreeSHAP analysis was used to interpret global and patient-specific contributions in the highest-AUC model. Postoperative abdominal distension occurred in 138 of 1,234 patients (11.18%). In the internal test set ( n = 247; 28 events), random forest (RF) had the highest AUC (0.775, 95% CI 0.650–0.880), followed by GB (0.746, 95% CI 0.627–0.854), SVM (0.737, 95% CI 0.635–0.831), and LR (0.723, 95% CI 0.616–0.821). The AUC difference between RF and GB was statistically significant ( P = 0.039); all other pairwise comparisons were not significant. TreeSHAP ranked PCA (mean absolute SHAP value, 0.041), surgical drainage tube placement (0.024), and operative time (OT)(0.023) as the leading global contributors. PCA use and drainage tube placement generally shifted predictions toward higher PAD risk, whereas higher postoperative ADL values generally shifted predictions toward lower risk. The models showed moderate internal discrimination but materially different sensitivity–specificity trade-offs. Permutation importance and TreeSHAP consistently identified PCA and drainage tube placement as leading random-forest predictors, while TreeSHAP additionally showed the direction and patient-specific magnitude of each contribution. These model explanations are predictive rather than causal. Prospective temporal validation, external validation, and threshold selection are required before bedside use.