Aug 2026· Frontiers in Medicine· Vol 13· 0 citations· 33 references
TL;DR
The machine learning model developed in this study demonstrates strong capability in stratifying anesthetic risk for patients with lumbar spinal stenosis, providing valuable reference for selecting surgical and anesthetic approaches.
Abstract
Patients with lumbar spinal stenosis are typically elderly with multiple comorbidities, necessitating accurate preoperative anesthetic risk assessment. The American Society of Anesthesiologists (ASA) classification quantifies functional reserve and disease burden, serving as a widely used tool for risk stratification. However, ASA classification is often influenced by subjective factors including physician experience and varies among clinicians with different seniority, while the assessment process remains time-consuming. This study aimed to develop an automated model for anesthetic risk stratification and evaluate its performance, with the goal of providing decision support for surgical and anesthetic management in this patient population.
Clinical data of 600 patients with lumbar spinal stenosis were collected and randomly divided into training (
n
= 480) and internal validation (
n
= 120) sets. An additional 100 patients from another tertiary hospital formed an external validation set. The model was validated and hyperparameter-tuned using k-fold cross-validation. Model performance, including overall classification and high-risk identification, was evaluated using accuracy, macro-average precision, macro-average recall, macro-average F1 score, weighted Kappa, linear weighted accuracy, positive predictive value, and negative predictive value.
In the internal validation set, the accuracy, macro-average precision, macro-average recall, macro-average F1 score, weighted Kappa coefficient and linear weighted accuracy of the model are 0.97, 0.96, 0.95, 0.96, 0.93 and 0.98 respectively, while in the external validation set, they are 0.97, 0.97, 0.94, 0.96, 0.93 and 0.99 respectively. The confusion matrix heatmap shows that the error is mainly concentrated between adjacent classes, and there is no cross-class misjudgment. In the internal validation set, the model's accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and Kappa coefficient for identifying high-risk patients were 0.98, 0.95, 0.98, 0.91, 0.99, and 0.92, respectively. In the external validation set, these values were 0.98, 0.93, 0.99, 0.93, 0.99, and 0.92, respectively.
The machine learning model developed in this study demonstrates strong capability in stratifying anesthetic risk for patients with lumbar spinal stenosis, providing valuable reference for selecting surgical and anesthetic approaches.
RF-based models can accurately and equitably predict perioperative complications in diverse spine surgery contexts, supporting personalized counseling, targeted monitoring, and optimized resource allocation.
Andrea Campagner, Francesco Langella, P. Bellosta-López et al.· European spine journal· 0 citations
Early identification of patients at risk of postoperative abdominal distension after spinal surgery remains clinically difficult. Machine learning may improve risk prediction by integrating routinely available perioperative information.
This study aimed to develop and internally validate an interpretable machine learning-based predictive model for postoperative abdominal distension after spinal surgery under general anesthesia and to identify clinically relevant predictors for early nursing risk stratification.
This retrospective single-center prediction model study included 1,234 patients who underwent spinal surgery under general anesthesia. Seventeen routinely recorded perioperative and early postoperative variables were evaluated, including postoperative ADL category(ADL), PCA, and postoperative antibiotic use (yes/no). The data were split using stratified sampling into training (80%) and internal test (20%) sets. Missing-value imputation, scaling, five-fold cross-validated grid search, and model fitting were performed within the training workflow. Logistic regression(LR), RF, support vector machine(SVM), and gradient boosting (GB)models were evaluated using discrimination, calibration, Brier score, decision curve analysis, and permutation importance. Exact TreeSHAP analysis was used to interpret global and patient-specific contributions in the highest-AUC model.
Postoperative abdominal distension occurred in 138 of 1,234 patients (11.18%). In the internal test set (
n
= 247; 28 events), random forest (RF) had the highest AUC (0.775, 95% CI 0.650–0.880), followed by GB (0.746, 95% CI 0.627–0.854), SVM (0.737, 95% CI 0.635–0.831), and LR (0.723, 95% CI 0.616–0.821). The AUC difference between RF and GB was statistically significant (
P
= 0.039); all other pairwise comparisons were not significant. TreeSHAP ranked PCA (mean absolute SHAP value, 0.041), surgical drainage tube placement (0.024), and operative time (OT)(0.023) as the leading global contributors. PCA use and drainage tube placement generally shifted predictions toward higher PAD risk, whereas higher postoperative ADL values generally shifted predictions toward lower risk.
The models showed moderate internal discrimination but materially different sensitivity–specificity trade-offs. Permutation importance and TreeSHAP consistently identified PCA and drainage tube placement as leading random-forest predictors, while TreeSHAP additionally showed the direction and patient-specific magnitude of each contribution. These model explanations are predictive rather than causal. Prospective temporal validation, external validation, and threshold selection are required before bedside use.
Unknown authors· Frontiers in Medicine· 0 citations
An interpretable gradient-boosting model may support risk-stratified perioperative assessment for elderly patients undergoing abdominal surgery and Prospective multicenter validation is required before routine clinical implementation.
Qiang Zhong, Guiming Huang, Wen Zhou et al.· Frontiers in Surgery· 0 citations
Background Early identification of patients at risk of inadequate postoperative analgesia remains a major challenge in perioperative medicine. Although machine learning approaches have shown potential for clinical outcome prediction, many existing models primarily emphasize discrimination performance while insufficiently addressing reliability, interpretability, and clinical applicability. This study aimed to develop and comprehensively evaluate an explainable artificial intelligence (XAI) framework for postoperative analgesia risk stratification and clinician-oriented decision support following abdominal surgery. Methods A retrospective cohort of 202 adult patients undergoing elective abdominal surgery and receiving postoperative analgesia with oxycodone or tilidine was analyzed. After exclusion of five patients with incomplete variables required for feature construction, 197 patient-level records were included in machine learning model development. Ten supervised learning algorithms were trained using demographic characteristics, surgical factors, analgesic information, and early postoperative pain trajectory features. Model performance was evaluated through stratified five-fold cross-validation using discrimination metrics, precision-recall analysis, calibration assessment, learning-curve analysis, and decision curve analysis (DCA). Model interpretability was assessed using SHAP-based feature attribution and cross-model explanation consistency analysis. Propensity score matching (PSM) was performed to evaluate the association between analgesic selection and postoperative outcomes after adjustment for measured baseline differences. A large language model (LLM) was integrated as a post hoc interpretation layer to translate model outputs into clinician-readable explanations. Results The evaluated models demonstrated moderate-to-good predictive performance, with validation AUC values ranging from 0.707 to 0.798. Early postoperative pain trajectory variables, particularly VAS-derived features, consistently represented the strongest predictors across different algorithms. Calibration analysis revealed differences in probability reliability that were not captured by discrimination metrics alone, while DCA demonstrated potential clinical utility of the best-performing models within clinically relevant decision thresholds. SHAP-based analyses showed stable feature attribution patterns across algorithms, supporting the robustness of identified predictors. After adjustment using PSM, analgesic type remained associated with postoperative analgesia outcomes, suggesting potential differences in real-world analgesic effectiveness while acknowledging the observational study design. Conclusions This study presents an integrated XAI framework combining predictive modeling, calibration assessment, decision-curve evaluation, explainability analysis, and language-based interpretation for postoperative analgesia risk stratification. The results highlight the importance of evaluating both predictive accuracy and clinical reliability when developing AI-assisted decision-support systems. Further multicenter prospective studies incorporating larger datasets and real-time clinical variables are required to validate model generalizability and evaluate implementation in perioperative care.
M. Wang, Jiao-Min Jian, Jun Wu· Frontiers in Digital Health· 0 citations
Background Pulmonary Embolism (PE) is a condition that results in significant mortality and morbidity, particularly among the elderly. The aim of our study is to utilize machine learning (ML) algorithms specifically tailored for this population, thereby enhancing the accuracy of risk assessment. Methods Elderly patients with PE who were included between January 1, 2013, and December 31, 2023, were divided into two groups: training and validation sets. A total of five ML models, including decision tree, random forest (RF), extreme gradient boosting, support vector machine, and k-nearest neighbors, were developed to predict in-hospital mortality in these patients. The model demonstrating the best diagnostic performance was selected. Ultimately, the ML models were internally validated to assess their diagnostic performance using receiver operating characteristic analysis. Results The analysis included 250 patients, with 174 assigned to the training set and 76 to the validation set. The ML models were developed using nine clinical features: age, lactate levels, blood urea nitrogen, arterial oxygen partial pressure, red blood cell count, diastolic blood pressure, vasopressor use, acute kidney injury, and the requirement for continuous renal replacement therapy. Of the five evaluated ML models, the RF model demonstrated the best performance, attaining the highest area under the curve values of 0.950 for the training set and 0.835 for the validation set. Conclusion ML-based models exhibit strong predictive capabilities for identifying elderly patients with PE who are at risk of hospital mortality. These algorithms can aid physicians in the early detection of high-risk patients, facilitating timely and appropriate preventive interventions. Future prospective studies comparing ML models with established PE-specific risk stratification tools are warranted.
Tao Chu, Ling Ji, Ding-yu Tan et al.· Frontiers in Medicine· 0 citations
BACKGROUND
Postoperative delirium is a common and serious complication after general anesthesia; its accurate prediction remains a substantial challenge in perioperative medicine. Existing models primarily rely on clinical variables and may have limited predictive accuracy. This study aimed to evaluate the added value of heart rate variability parameters in predicting postoperative delirium and construct an interpretable multimodal predictive model.
METHODS
In this prospective observational study, 1418 patients undergoing general anesthesia were included. Seventy-three features, including electrocardiogram abnormalities and heart rate variability time-, frequency-, and nonlinear-domain indicators, were extracted from electrocardiogram data. Postoperative delirium was assessed using the Chinese version of the 3-Minute Diagnostic Interview for Delirium within 3 days postoperatively. Feature selection was conducted by combining least absolute shrinkage and selection operator (LASSO) regression, the Boruta algorithm, and random forests, and 10 machine learning models were developed. Model performance was evaluated through receiver operating characteristic curves and decision curve analysis, with interpretability assessed via Shapley additive explanations. Clinical prediction tools were derived from key features. We used an external validation set to further evaluate the generalization ability of the models.
RESULTS
Postoperative delirium occurred in 255 (18%) patients. Seventeen key predictors were identified in total. The combined clinical-electrocardiogram-heart rate variability model demonstrated the highest predictive performance (area under the curve = 0.728), outperforming clinical-only (area under the curve = 0.673) and electrocardiogram-only models (area under the curve = 0.679). Logistic regression showed the highest discrimination. In the external validation set, the model maintained robust performance with an area under the curve value of 0.836. Shapley additive explanations highlighted seven core predictors: atrial or ventricular arrhythmia, operative time, ST-segment abnormalities, age, American Society of Anesthesiologists classification, heart rate variability entropy, and overall electrocardiogram abnormalities. A nomogram and online platform enabled personalized risk assessment.
CONCLUSIONS
Our results indicate that integrating heart rate variability with clinical and electrocardiogram features significantly enhances the personalized predictive efficacy of postoperative delirium.
Yuling Tang, Yuanhui Liu, Jiayi Tang et al.· Anesthesia and Analgesia· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.