Skip to content
Conference

Machine Learning-based Predictive Analytics for Early Detection of Chronic Diseases

Jul 2026 · 2026 7th International Conference on Smart Systems and Inventive Technology (ICSSIT) · pp. 2060-2066 · 0 citations · 21 references

Abstract

Chronic diseases remain a major cause of mortality and long-term disability, and proactive identification of high-risk individuals is difficult because early clinical changes are often subtle, incomplete, and distributed across heterogeneous hospital records. This study proposes a machine-learning-based predictive analytics framework for early detection of chronic disease risk using electronic health records, laboratory profiles, demographic factors, medication history, and derived clinical indicators. The study used 48,320 adult patient records collected from four tertiary hospitals between 2014 and 2023, with leakage-controlled preprocessing, multistage missing-data handling, correlation and SHAP-assisted feature selection, and stratified model development. Logistic Regression, Random Forest, XGBoost, Multilayer Perceptron, and TabNet were evaluated against clinical risk-score baselines using AUROC, PR-AUC, F1-score, recall, calibration, Brier score, and stability under missingness and imbalance. XGBoost achieved the strongest internal performance with AUROC of 0.942, PR-AUC of 0.901, F1-score of 0.901, and Brier score of 0.108. External validation on 12,950 patients from an unseen hospital produced AUROC of 0.931, confirming limited performance degradation and improved generalizability. SHAP analysis identified creatinine, HbA1c, age, systolic blood pressure, and triglycerides as dominant contributors, supporting clinically interpretable early-risk alerts for preventive care.

View source

Similar papers

Open access Jul 2026

Machine Learning-Based Early Prediction of Hospital Readmission Risk Among Chronic Disease Patients Using Electronic Health Records: A Comparative Study of Ensemble Learning Models

Findings indicate that XGBoost effectively identifies patients at high risk of early hospital readmission and can serve as a reliable predictive tool for clinical decision support.

Md Yassir Mottalib, Eklachur Rahman Bhuiyan, Anwar Hossain et al. · 0 citations
Open access Aug 2026

A Comparative Evaluation of Machine Learning Algorithms for Diabetes Risk Prediction

Diabetes mellitus is a chronic metabolic disorder whose global prevalence continues to rise, creating an urgent need for scalable, low-cost tools for early risk identification. This study evaluates the effectiveness of five machine learning algorithms (Random Forest, XGBoost, Support Vector Machine [SVM], CatBoost, and TabNet) for predicting diabetes risk from routinely available clinical and lifestyle variables. Using the Pima Indians Diabetes dataset, a preprocessing pipeline was applied that included median imputation of physiologically implausible zero values, standardized (Z-score) feature scaling, Boruta-based feature selection, and class-imbalance handling through class-weight adjustment and the Synthetic Minority Oversampling Technique (SMOTE). Models were trained on an 80/20 train-test split and assessed using accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (ROC-AUC). Random Forest achieved the strongest overall performance (accuracy 0.753; F1-score 0.689; ROC-AUC 0.810), followed by CatBoost, SVM, and XGBoost, whereas TabNet performed worst with very low recall for the diabetic class. The best-performing model (Random Forest) was deployed in a lightweight Flask web application that returns a probability-based diabetes risk assessment, categorising each prediction as low, moderate, or high risk together with a tailored recommendation. The findings confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained settings. Key limitations include dataset homogeneity, residual class imbalance, and limited feature coverage.

T. Olayinka · 0 citations
Open access Aug 2026

Clinical Biomarker-Based Prediction of Chronic Kidney Disease Using Explainable Machine Learning

The results show how a combination of explainable ML and accessible clinical biomarkers can offer a precise, transparent, and clinically interpretable framework for early CKD diagnosis, risk stratification, and informed clinical decision making.

M. Khuntia, Hariballav Mahapatra, N. Lodha · 0 citations
Open access Aug 2026

PREDICTING HEART DISEASE RISK FROM CLINICAL VARIABLES: A GENDER-SPECIFIC MACHINE LEARNING ANALYSIS AMONG HIGH-CHOLESTEROL PATIENTS

Male sex was a statistically significant independent predictor of heart disease after controlling for other clinical variables and the findings support sex-specific screening and preventive strategies for high-cholesterol male patients and demonstrate the value of interpretable machine learning models for clinical decision support.

Taiwo Samson Adeyemo · 0 citations
Open access Jul 2026

Stroke prediction using Ensemble Learning

The proposed model employs Ensemble Learning techniques, which combine multiple machine learning algorithms to improve prediction accuracy and robustness, and is capable of identifying complex patterns in medical data and classifying patients into stroke-risk categories with high efficiency.

Bhagyashri Patil, Priyadarshini C Patil, Soumya M A et al. · 0 citations
Conference Jul 2026

AI-Powered Predictive Model for Early Detection of Disease Progression using Patient Health Data

Routine clinical metrics may miss subtle physiological variations that occur before clinically evident disease progression, delaying preventive intervention. This study presents an AI-powered predictive model trained on 4,567 de-identified longitudinal patient records containing 36 structured clinical variables, 14 laboratory biomarkers, demographic descriptors, medication history, and three years of follow-up. Disease progression was labelled using a predefined composite endpoint combining sustained biomarker deterioration, clinically documented worsening, treatment escalation, or disease-related hospitalization within the follow-up period. Missing values were managed using temporally constrained forward-backward interpolation with missingness indicators, class imbalance was examined using progression and non-progression distributions, and privacy was maintained through de-identification and controlled data handling. A hybrid temporal convolutional encoder and gated recurrent prediction module captured short-term fluctuations and long-range trends. The model was trained on 3,214 records and tested on 1,353 records, achieving 94.27% accuracy, 0.962 AUC, and 0.943 F1-score, while reducing false negatives compared with conventional clinical scoring and sequential baselines. Generalization was assessed through patient-level hold-out testing, temporally separated validation, and stratified five-fold cross-validation because a fully independent external cohort was not available. These findings support the potential of the model for early-warning prediction and clinically timely intervention.

P. A. Prakash, Aakila Fathima S, Divyadharshini S et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.