Jul 2026· International Scientific Journal of Engineering and Management· Vol 05, pp. 1-9· 0 citations
TL;DR
The results demonstrate that ML models can effectively identify individuals at high risk of hypertension, offering a valuable tool for early intervention and personalized healthcare and underscores the potential of artificial intelligence in supporting public health efforts and enhancing clinical decision-making.
Abstract
Hypertension, commonly known as high blood pressure, is a major risk factor for cardiovascular diseases and premature mortality worldwide. Early detection and prevention are critical in reducing its health impact. This study explores the application of machine learning (ML) techniques to predict the likelihood of hypertension in individuals using clinical and demographic data. A variety of supervised learning algorithms, including Logistic Regression, Random Forest, Support Vector Machines, and Gradient Boosting, were evaluated for their predictive performance [1]. The dataset was preprocessed through feature selection, normalization, and handling of missing values to improve model accuracy.[2] Performance metrics such as accuracy, precision, recall, F1-score, and AUC-ROC were used to assess the models [4]. The results demonstrate that ML models can effectively identify individuals at high risk of hypertension, offering a valuable tool for early intervention and personalized healthcare [5]. This approach underscores the potential of artificial intelligence in supporting public health efforts and enhancing clinical decision-making.
Key words: Logistic Regression, Random Forest, Support Vector Machines, and Gradient Boosting.
There is an urgent need for explainable, clinically validated and standardised ML frameworks to translate predictive models into routine healthcare practice and improve early detection of cardiovascular disease.
Hanna Rasheed, Arya.K.R Arya.K.R, Ashida.K.A Ashida.K.A· International Journal of Tec...· 0 citations
Hypertension is one of the most important modifiable risk factors for Cardiovascular Disease (CVD), yet identifying which hypertensive patients are at higher risk remains challenging in clinical practice. This study developed and evaluated three machine-learning models: logistic regression, random forest, and Gradient Boosting for CVD risk prediction in a cohort of 23,543 hypertensive patients drawn from a 70,000 patient cardiovascular dataset. After preprocessing, feature engineering, SMOTE-based class balancing, and hyperparameter tuning via randomized search, model performance was assessed on a held-out test set and validated using 5-fold stratified cross-validation with SMOTE correctly nested inside each fold to avoid data leakage. On the test set, tuned Gradient Boosting model achieved the highest accuracy (78.59%) and AUC-ROC (0.6681), outperforming Logistic Regression (0.6633) and Random Forest (0.6508). cross-validation provided a slightly different perspective: Logistic Regression’s mean AUC-ROC (0.6628) edged out Gradient Boosting (0,6609) and Random Forest (0.6383), SHAP analysis on the Gradient Boosting model identified systolic blood pressure, age, and height as the strongest predictors, with height rivaling systolic blood pressure and surpassing BMI a notable difference from Random Forest’s feature importance ranking. Lifestyle factors (smoking, alcohol, physical activity) contributed minimally. These findings highlight blood pressure and body size measures as the dominant clinical signals in this dataset, while demonstrating the potential of an explainable machine-learning model based on routinely collected clinical data to support cardiovascular risk stratification and clinical decision-making in hypertensive patients, despite their moderate discriminative performance.
C. M. Anyanwu, J. C. Onyianta, Ogechi Gift Onyedi et al.· Nature Journal of Emerging S...· 0 citations
Despite being one of the major modifiable risk factors for cardiovascular disease, stroke and premature death globally, a significant proportion of those affected are not detected or managed well. This paper presents the design and development of a real-time predictive system for early detection of hypertension risk by continuously streamed physiological parameters such as BP, HR, Peripheral oxygen saturation (SpO2), and body temperature. The engineered features such as pulse pressure, mean arterial pressure (MAP), and rate-pressure product (RPP) are calculated from raw sensor readings and then used to train supervised machine learning classifiers namely logistic regression, random forest, support vector machine (SVM), and gradient boosting. A labelled set of 12,000 physiological signals were used to train and validate models through stratified 5-fold cross-validation and hyper-parameter optimization via grid-search. The models (LDA, SVM, NB, RF) performed well, with the RF model chosen as the deployed model, yielding the highest test accuracy of 99.8%, precision of 100%, recall of 99.2%, F1-score of 0.996 and ROC-AUC of 1.000. The trained model was embedded in a real-time dashboard that feeds streaming sensor data into the model for real-time risk calculation and alerts the clinician when the risk for hypertension is greater than a preset limit, with a retraining loop using feedback to improve the model. The results show that the use of cheap IoT sensing in combination with light weight machine learning for continuous, real-time hypertension screening and early-warning support is feasible, both in clinical settings and in home monitoring.
Dhanna Ram, Somil Jain, Rajesh Yadav et al.· International Research Journ...· 0 citations
The proposed model employs Ensemble Learning techniques, which combine multiple machine learning algorithms to improve prediction accuracy and robustness, and is capable of identifying complex patterns in medical data and classifying patients into stroke-risk categories with high efficiency.
Bhagyashri Patil, Priyadarshini C Patil, Soumya M A et al.· International journal of com...· 0 citations
Diabetes mellitus is a chronic metabolic disorder whose global prevalence continues to rise, creating an urgent need for scalable, low-cost tools for early risk identification. This study evaluates the effectiveness of five machine learning algorithms (Random Forest, XGBoost, Support Vector Machine [SVM], CatBoost, and TabNet) for predicting diabetes risk from routinely available clinical and lifestyle variables. Using the Pima Indians Diabetes dataset, a preprocessing pipeline was applied that included median imputation of physiologically implausible zero values, standardized (Z-score) feature scaling, Boruta-based feature selection, and class-imbalance handling through class-weight adjustment and the Synthetic Minority Oversampling Technique (SMOTE). Models were trained on an 80/20 train-test split and assessed using accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (ROC-AUC). Random Forest achieved the strongest overall performance (accuracy 0.753; F1-score 0.689; ROC-AUC 0.810), followed by CatBoost, SVM, and XGBoost, whereas TabNet performed worst with very low recall for the diabetic class. The best-performing model (Random Forest) was deployed in a lightweight Flask web application that returns a probability-based diabetes risk assessment, categorising each prediction as low, moderate, or high risk together with a tailored recommendation. The findings confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained settings. Key limitations include dataset homogeneity, residual class imbalance, and limited feature coverage.
T. Olayinka· FUDMA Journal of Sciences· 0 citations
Experimental results demonstrate that Machine Learning techniques can effectively predict disease occurrence with high accuracy, thereby assisting healthcare professionals in early diagnosis and treatment planning.
Sunidhi, Rahul, Parmod Kumar et al.· International Journal for Re...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.