Machine Learning-Based Early Prediction of Hospital Readmission Risk Among Chronic Disease Patients Using Electronic Health Records: A Comparative Study of Ensemble Learning Models
Jul 2026· Frontline Medical Sciences and Pharmaceutical Journal· 0 citations
TL;DR
Findings indicate that XGBoost effectively identifies patients at high risk of early hospital readmission and can serve as a reliable predictive tool for clinical decision support.
Abstract
Hospital readmission among patients with chronic diseases remains a major challenge for healthcare systems due to its association with poor patient outcomes and increased healthcare costs. This study proposes a machine learning-based framework for the early prediction of 30-day hospital readmission risk using the publicly available Diabetes 130-US Hospitals dataset from the UCI Machine Learning Repository. A comprehensive preprocessing pipeline, feature engineering, and feature selection techniques were employed to improve data quality and predictive performance. Eight supervised machine learning algorithms, including Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, LightGBM, CatBoost, Multilayer Perceptron, and XGBoost, were developed and comparatively evaluated. Model performance was assessed using accuracy, precision, recall, F1-score, specificity, and the area under the receiver operating characteristic curve (AUC-ROC). The experimental results demonstrated that ensemble learning models consistently outperformed conventional machine learning approaches. Among all evaluated models, XGBoost achieved the best performance, attaining 92.16% accuracy, 0.92 precision, 0.91 recall, 0.91 F1-score, 0.95 specificity, and an AUC-ROC of 0.972. These findings indicate that XGBoost effectively identifies patients at high risk of early hospital readmission and can serve as a reliable predictive tool for clinical decision support. The proposed framework has strong potential for integration with Electronic Health Record systems to facilitate early intervention, improve patient outcomes, reduce preventable readmissions, and support value-based healthcare delivery.
Hospital readmission is a serious problem in healthcare systems as it leads to higher treatment costs, resource use and patient morbidity. Identifying patients who are likely to be readmitted can help guide prompt action by clinical staff and enhance health outcomes. In this study, a comprehensive benchmarking analysis of supervised machine learning classifiers for diabetic patients is presented. The data pre-processing steps comprised of missing value handling, drop of identifier attributes, categorical feature encoding, and binary target generation. In order to overcome class imbalance, the Synthetic Minority Over-sampling Technique (SMOTE) was used only for the training set, and the original test set was kept imbalanced to evaluate the algorithm in an unbiased manner. The performance of eight machine learning classifiers Logistic Regression, Decision Tree, Random Forest, Extra Trees, AdaBoost, Gradient Boosting, XGBoost and LightGBM was assessed. The experimental results show the superiority of ensemble learning approaches over the conventional classifiers in predicting hospital readmission. Extra Trees with 86.54% classification accuracy obtained the best ROC AUC 0.6079 among models and with the training time of 1.87 s, it shows a balance between the predictive power and computational efficiency. Random Forest had the highest precision (21.58%), AdaBoost had the highest recall (33.42%) and F1 score (0.1919), Logistic Regression had the highest balanced accuracy (53.35%) and XGBoost had the highest Matthews correlation coefficient (0.0759). All the findings suggest that no one classifier was superior in all the evaluation measures, which illustrates the need for multi-metric evaluation in designing predictive models for imbalanced healthcare datasets. This benchmark study offers an experimental framework and hands-on guidance on selecting appropriate machine learning models for hospital-readmission predictions in diabetes management.
A. Gupta, Arun Kumar Choudhary· Genetics and Molecular Resea...· 0 citations
Chronic diseases remain a major cause of mortality and long-term disability, and proactive identification of high-risk individuals is difficult because early clinical changes are often subtle, incomplete, and distributed across heterogeneous hospital records. This study proposes a machine-learning-based predictive analytics framework for early detection of chronic disease risk using electronic health records, laboratory profiles, demographic factors, medication history, and derived clinical indicators. The study used 48,320 adult patient records collected from four tertiary hospitals between 2014 and 2023, with leakage-controlled preprocessing, multistage missing-data handling, correlation and SHAP-assisted feature selection, and stratified model development. Logistic Regression, Random Forest, XGBoost, Multilayer Perceptron, and TabNet were evaluated against clinical risk-score baselines using AUROC, PR-AUC, F1-score, recall, calibration, Brier score, and stability under missingness and imbalance. XGBoost achieved the strongest internal performance with AUROC of 0.942, PR-AUC of 0.901, F1-score of 0.901, and Brier score of 0.108. External validation on 12,950 patients from an unseen hospital produced AUROC of 0.931, confirming limited performance degradation and improved generalizability. SHAP analysis identified creatinine, HbA1c, age, systolic blood pressure, and triglycerides as dominant contributors, supporting clinically interpretable early-risk alerts for preventive care.
P. A. Prakash, Mamtha C, Manishathri R et al.· 2026 7th International Confe...· 0 citations
Objective This study aimed to develop and validate a predictive model using machine learning to estimate mortality among intensive care unit patients. Methods The medical records of 874 sepsis patients hospitalized at the Affiliated Hospital of Chengde Medical University from 2021 to 2024 were retrospectively analyzed. Sepsis patients were randomly divided into training and validation sets in a 7:3 ratio. We constructed mortality prediction models for sepsis patients using machine learning algorithms, including extreme gradient boosting (XGBoost), logistic regression (LR), random forest (RF), adaptive boosting (AdaBoost), k-nearest neighbors (KNN), support vector machine (SVM), multilayer perceptron (MLP), and Gaussian naive Bayes (GNB). The predictive performance of the machine learning models was assessed using receiver operating characteristic (ROC) curves, calibration curves, and decision curve analysis (DCA). This study was reported in accordance with the RECORD (REporting of studies Conducted using Observational Routinely collected health Data) guidelines. Results A total of 874 patients with sepsis were included, of whom 338 died and 536 survived. Significant differences were observed between the mortality and survival groups with respect to platelet count (PLT), platelet distribution width (PDW), platelet distribution width to count ratio (PCR), mean platelet volume (MPV), monocyte count (Mono), albumin (Alb) level, total bilirubin (Tbil), alanine aminotransferase (ALT), aspartate aminotransferase (AST), lactate dehydrogenase (LDH), serum creatinine (Scr), blood urea nitrogen (BUN), fibrinogen (Fib) level, D-dimer level, lactate (Lac) level, respiratory system infection, gastrointestinal system infection, and APACHEII score. The ROC curve analysis showed that the areas under the curve (AUC) of the XGBoost, LR, RF, AdaBoost, KNN, SVM, MLP and GNB models in predicting the in-hospital mortality rate of sepsis patients were 0.954, 0.880, 0.951, 0.901, 0.859, 0.910, 0.836 and 0.854, respectively. Among the nine algorithms, the XGboost model performed significantly better than the others, achieving an accuracy of 0.879, a sensitivity of 0.734, and an F1 score of 0.821. The calibration curve demonstrated that the XGBoost model showed the best performance among the eight algorithms, with predictions closely matching the observed outcomes. Decision curve analysis indicated that the XGBoost model outperformed both extreme strategies. Conclusion Machine learning models provide a reliable approach for predicting in-hospital mortality in patients with sepsis. Among them, the XGboost model demonstrated the highest predictive performance, aiding clinicians in identifying high-risk patients with sepsis and implementing early interventions to reduce mortality.
Tingting Wang, Yi Sun, Mengna Zhang et al.· Journal of Inflammation Rese...· 0 citations
BACKGROUND
Coronary artery disease (CAD) is the leading cause of death globally and a major contributor to hospital readmission. This study aimed to predict 30-day mortality in patients hospitalized with acute and chronic CAD using a structured machine learning approach with data from multiple centers.
METHODS
We conducted a retrospective cohort study using patient data from the Taipei Medical University Clinical Research Database (TMUCRD). Multiple machine learning algorithms were employed to develop predictive models for 30-day mortality. Model performance was evaluated using a stratified fivefold cross-validation approach. Key performance metrics included the area under the curve (AUC), accuracy, sensitivity, specificity, negative predictive value (NPV), positive predictive value (PPV), and F1 score.
RESULTS
A total of 23,267 patients (mean age 64.9 years) were included, with 1215 deaths overall (5.2%): 570 (3.7%) in the internal cohort (n = 15,510) and 645 (8.3%) in the external validation cohort (n = 7757, Shuang Ho Hospital). XGBoost achieved the best performance for the overall and acute CAD cohorts (AUROC 0.845 and 0.820, respectively), while logistic regression performed best for chronic CAD (AUROC 0.766). Key predictive features included the Charlson Comorbidity Index, hemoglobin level, emergency room admission status, age, and creatinine level.
CONCLUSION
The use of a structured machine learning approach to predict 30-day mortality in patients with acute and chronic CAD demonstrated promising discriminative performance, providing valuable insights that could enhance personalized care and inform clinical decisions.
Septi Melisa, P. Phan, Sheng-Hsuan Chien et al.· International Journal of Car...· 0 citations
Cardiovascular disease remains one of the leading causes of mortality worldwide, necessitating the development of accurate and interpretable predictive systems that support early diagnosis and clinical decision-making. While numerous machine learning models have demonstrated promising predictive capabilities, many operate as black-box systems that provide limited transparency regarding how predictions are generated. This lack of interpretability presents significant challenges in healthcare environments where trust, accountability, and regulatory compliance are essential. This study presents a comparative evaluation of six supervised machine learning classifiers for heart disease prediction using the Cleveland Heart Disease Dataset. The evaluated models include Decision Tree, Logistic Regression, Support Vector Machine, K-Nearest Neighbours, Random Forest, and Gradient Boosting. A comprehensive machine learning pipeline comprising data preprocessing, feature selection, hyperparameter optimization, repeated stratified cross-validation, and performance evaluation was implemented. Explainable Artificial Intelligence (XAI) techniques based on SHAP were integrated to provide both global and local interpretability of model predictions. Experimental results demonstrate that the Decision Tree classifier achieved the highest overall performance, attaining an accuracy of 98.54%, precision of 98.21%, recall of 98.34%, and F1-score of 98.27%. Furthermore, SHAP-based analysis revealed that chest pain type, number of major vessels, exercise-induced angina, maximum heart rate, and ST depression were the most influential predictors of cardiovascular disease risk. The findings indicate that interpretable machine learning models can achieve predictive performance comparable to or exceeding more complex algorithms while maintaining transparency and clinical usability. The study contributes a reproducible framework for explainable cardiovascular disease prediction and demonstrates the feasibility of integrating interpretable machine learning models into clinical decision support systems. The proposed approach offers a foundation for trustworthy healthcare artificial intelligence applications that balance predictive accuracy with explainability.
Asoshi Paul Anule, Anagu Emmanuel John, Ogar Michael Oko· Middle East Journal of Appli...· 0 citations
Male sex was a statistically significant independent predictor of heart disease after controlling for other clinical variables and the findings support sex-specific screening and preventive strategies for high-cholesterol male patients and demonstrate the value of interpretable machine learning models for clinical decision support.
Taiwo Samson Adeyemo· GSC Advanced Research and Re...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.