Skip to content
Open access

PREDICTING HEART DISEASE RISK FROM CLINICAL VARIABLES: A GENDER-SPECIFIC MACHINE LEARNING ANALYSIS AMONG HIGH-CHOLESTEROL PATIENTS

Aug 2026 · GSC Advanced Research and Reviews · 0 citations

Abstract

Cardiovascular disease remains a major cause of mortality and economic burden in the United States. This study developed and evaluated supervised machine learning models to predict heart disease risk from routine clinical variables and tested whether male patients with high cholesterol have higher odds of heart disease than female patients. Using a publicly available clinical dataset of 918 patients, the Cross-Industry Standard Process for Data Mining (CRISP-DM) framework guided exploratory analysis, data cleaning, median imputation of irregular cholesterol values, and feature selection using principal component analysis (PCA) and SelectKBest. Logistic Regression, Support Vector Machine (SVM), and a soft voting ensemble were trained and tuned using grid search with stratified 5-fold cross-validation. The ensemble achieved the highest predictive performance (accuracy = 0.940, F1 = 0.950, receiver operating characteristic area under the curve [ROC-AUC] = 0.958). Logistic Regression achieved comparable performance (accuracy = 0.929, F1 = 0.940, ROC-AUC = 0.958) and was selected for hypothesis testing because of its interpretability. Among patients with cholesterol ≥240 mg/dL, male sex was a statistically significant independent predictor of heart disease after controlling for other clinical variables. The findings support sex-specific screening and preventive strategies for high-cholesterol male patients and demonstrate the value of interpretable machine learning models for clinical decision support. Larger externally validated datasets are needed to assess generalizability.

Read PDF