Skip to content
Preprint

Transforming Heart Disease Prediction with Advanced Machine Learning Techniques

Aug 2026 · 0 citations · 41 references
Computer Science

TL;DR

The research work concludes that ML models, when properly tuned and validated, can significantly assist in the early diagnosis of heart disease, offering critical support for clinical decision-making.

Abstract

Heart disease remains the leading cause of mortality globally, necessitating early and accurate detection to improve patient outcomes. This research focuses on the predictive analysis of heart disease using machine learning (ML) techniques, comparing the performance of multiple classifiers to identify the most accurate and least error-prone method. Two datasets from UCI and Kaggle repositories were utilized, each containing 14 attributes related to heart health indicators. Techniques including J48, Naive Bayes, Logistic Regression, Simple Cart, Bagging, Decision Stump, AdaBoost, Artificial Neural Networks, and Support Vector Machine (SVM) were applied. Evaluation metrics such as Mean Absolute Error (MAE), Relative Absolute Error (RAE), accuracy, precision, recall, and F-measure were used for performance comparison. Results revealed that SVM achieved the highest performance on the UCI dataset, while Simple Cart performed best on the Kaggle dataset, offering the highest accuracy and lowest error rates. The research work concludes that ML models, when properly tuned and validated, can significantly assist in the early diagnosis of heart disease, offering critical support for clinical decision-making. Future work may involve hybrid approaches and the use of more recent datasets to further improve prediction accuracy.

View source

Similar papers

Open access Aug 2026

USING OPTIMAL MACHINE LEARNING ALGORITHMS TO PREDICT HEART FAILURE PATIENT CLASSIFICATION

Heart failure (HF) remains one of the leading causes of mortality worldwide, making early prediction and diagnosis essential for improving patient survival and reducing healthcare costs. Machine learning (ML) techniques have demonstrated considerable potential in assisting clinicians with accurate disease prediction. However, most heart failure datasets suffer from class imbalance, which negatively affects classification performance, particularly for minority class patients. This paper presents an optimized Extreme Gradient Boosting (XGBoost) model integrated with the Synthetic Minority Over-sampling Technique (SMOTE) for heart failure patient classification. Initially, missing values, outliers, and redundant attributes are removed through preprocessing. SMOTE is then applied to balance the dataset by generating synthetic minority samples. Hyperparameter optimization using Grid Search with Stratified Cross-Validation identifies the optimal XGBoost parameters. The proposed framework is evaluated using Accuracy, Precision, Recall, F1-score, ROC-AUC, and Matthews Correlation Coefficient (MCC). Experimental results demonstrate that the optimized XGBoost-SMOTE model significantly outperforms traditional machine learning algorithms including Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, KNearest Neighbors, AdaBoost, and baseline XGBoost. The proposed approach achieves an accuracy of 98.21%, precision of 97.94%, recall of 98.47%, F1-score of 98.20%, and ROC-AUC of 99.10%, indicating superior predictive capability for heart failure diagnosis. These findings suggest that integrating SMOTE with optimized XGBoost provides an effective decision-support tool for clinical risk assessment. Similar findings have been reported in prior studies evaluating XGBoost with SMOTE-based preprocessing for heart failure prediction.

B. Naveen, N. Rao · 0 citations
Open access Jul 2026

Performance Comparison of Tree-Based Models for Heart Disease Prediction Using Feature Selection and SMOTE

Heart disease remains the leading cause of mortality worldwide, highlighting the need for accurate early prediction models. This study proposes a machine learning framework for heart disease prediction using the BRFSS 2015 Heart Disease Health Indicators dataset, which contains 253,680 records and 22 attributes. The proposed approach integrates Synthetic Minority Oversampling Technique (SMOTE) for class imbalance handling, mutual information-based SelectKBest feature selection (k = 15), and three tree-based classifiers: Decision Tree, Random Forest, and XGBoost. A leakage-free preprocessing pipeline was implemented to ensure that SMOTE was applied only to the training data, and classification threshold optimization was performed to improve minority class detection. Model performance was evaluated using Accuracy, Precision, Recall, F1-score, and ROC-AUC metrics. Experimental results show that XGBoost achieved the best performance with a cross-validation ROC-AUC of 0.9815 and a test ROC-AUC of 0.8444 at an optimized threshold of 0.20. The findings demonstrate that the proposed integration of oversampling, feature selection, and threshold optimization can improve predictive performance for imbalanced cardiovascular risk data, providing a practical foundation for machine learning–based decision support in early heart disease risk screening.  

Santi Santi, Ema Utami · 0 citations
Open access Jul 2026

Heart Disease Prediction Using Logistic Regression and K-Nearest Neighbor: A Comparative Study of Classification Algorithm Performance

The findings suggest that Logistic Regression is more suitable as a decision-support model for early heart disease screening due to its higher sensitivity, accuracy, and specificity.

Yan Risa, Aspi Sururi, M. Asadullah et al. · 0 citations
Open access Aug 2026

An Advanced Ensemble Framework for Robust Heart Disease Detection and Classification

Cardiovascular Disease (CVD) remains one of the leading causes of mortality worldwide, emphasizing the need for accurate and early diagnostic solutions. Recent advances in Machine Learning (ML) and Deep Learning (DL) have shown significant potential to support clinical decision-making through data-driven prediction models. This study presents a robust Ensemble Learning (EL) framework for the prediction and classification of CVD by integrating multiple ML algorithms with a DL component. Specifically, an Artificial Neural Network (ANN) is employed as a feature extraction layer prior to ensemble aggregation using techniques such as Random Forest, XGBoost, and LightGBM. The proposed approach is evaluated using accuracy, precision, recall, F1-score, and AUC-ROC. Experimental results on a benchmark dataset demonstrate that the model achieves a high accuracy of 98.8%, outperforming individual classifiers and existing approaches. The integration of ANN-based feature extraction enhances model generalization and reduces prediction error. These findings highlight the effectiveness of the proposed framework for early heart disease detection and clinical decision support.

El Haddad Khadija, A. Bekkari, W. Bouarifi et al. · 0 citations
Review Open access Aug 2026

Artificial Intelligence for Early Heart Disease Prediction: A Review of Machine Learning Techniques

There is an urgent need for explainable, clinically validated and standardised ML frameworks to translate predictive models into routine healthcare practice and improve early detection of cardiovascular disease.

Hanna Rasheed, Arya.K.R Arya.K.R, Ashida.K.A Ashida.K.A · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.