Skip to content
Open access

Conceptualising heart disease prediction through a unified framework combining clinical theory and machine learning models

Aug 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 34 references
Computer Science

TL;DR

The results show that it is feasible to make better predictions and gain valuable insights by merging these two types of data and that integrating unstructured data allows for a more holistic view of patient health, leading to earlier detection, personalized interventions, and improved decision-making in clinical settings.

Abstract

A hybrid framework for a comprehensive evaluation of machine learning and transformer-based models for cardiovascular risk prediction by integrating structured clinical data with unstructured clinical narratives is included in this study. The structured data includes demographic and diagnostic variables like patient information and test results, while unstructured data is derived from electronic health records such as clinical notes. Through preprocessing, feature engineering, and semantic embedding using transformer-based models, the system leverages the complementary strengths of both data types. By analysing and processing this unstructured information, this project aims to improve the predictions for heart disease. The results show that it is feasible to make better predictions and gain valuable insights by merging these two types of data. Such analysis is essential because structured data alone often overlooks fine-grained clinical indicators found in narrative texts. Integrating unstructured data allows for a more holistic view of patient health, leading to earlier detection, personalized interventions, and improved decision-making in clinical settings.

Read PDF

Similar papers

Open access Aug 2026

Can GPT Be Used as an Alternative Prediction Model to Traditional Machine Learning and Neural Networks on Low-Volume Clinical Data?

The proposed GPT2-based table-to-text framework provides a practical and clinically interpretable approach for disease prediction from limited structured healthcare data and demonstrates strong potential for early risk detection, transparent clinical decision support, and reliable deployment in real-world low-resource healthcare environments.

S. Bin Akter, S. Akter, D. Eisenberg et al. · 0 citations
Conference Open access 2026

Natural Language Processing for Prediction of Chronic Diseases from Electronic Health Records

This approach combines semantic understanding of clinical narratives with structural modeling of patient-disease-treatment relationships and successfully validates synthetic EHR data utility for privacy-preserving healthcare AI development while addressing critical requirements necessary for clinical decision support system.

U. Luke, P. Asuquo, Victor Anaga et al. · 0 citations
Open access Aug 2026

Enhancing Cardiovascular Disease Diagnosis With Data-Driven Predictive Systems

The proposed approach uses a Quantum Neural Network for machine learning for machine learning in an intelligent Cardiovascular Disease (CVD) prediction system that has the highest sensitivity and specificity in the current literature, matching exact expert opinions.

Hutashani B. Rayate, Mangesh D. Nikose, Prakash G. Burade · 0 citations
Open access Jul 2026

Enhancing Cardiovascular Risk Prediction with Explainable AI using Clinical Data

Diabetes is a major risk factor for the development of cardiovascular issues which contribute to cardiovascular disease (CVD) being a leading cause of mortality worldwide. However, traditional machine learning methods are not widely adopted in healthcare systems because they lack interpretability, which is important for early and accurate CVD risk prediction and for ruling out effective clinical intervention. In this research, a hybrid architecture is proposed that incorporates diabetes related datasets as well as explainable artificial intelligence (XAI) methodologies that could improve the prediction power and transparency of the models. The proposed approach combines different datasets at the level of features and includes rigorous data pre-processing to detect metabolic and cardiovascular risk factors. Some of the significant clinical parameters are age, BMI, glucose, cholesterol, and blood pressure. These are standardized to create a single dataset which may be utilized for predictive modelling. The employment of two XAI approaches, SHAP (SHapley Additive Explanations) with tree-based ensemble models and integrated gradients with transformer based topologies, ensures both performance and interpretability. The technique improves confidence and usefulness in clinical settings by offering accurate predictions and explanations for the model’s judgments that are relevant to the circumstance. It is also utilized for visual investigation of clinical correlations of diabetes and cardiovascular disease and identify crucial risk variables. The results suggest that merging explainability approaches with powerful machine learning can considerably boost early identification and risk assessment. The proposed approach contributes to enhanced healthcare decision-making, offering a scalable, interpretable and dependable solution for cardiovascular disease prediction.

K. Deepthi, P. Bhargavi · 0 citations
Open access Jul 2026

An interpretable machine learning framework for early-stage diabetes mellitus prediction using comparative classification models and SHAP

A machine learning-based framework enhanced with explainability is introduced, built around a structured data preparation process that handles categorical encoding, numerical scaling, and minority class oversampling through the SMOTE technique, positioning it as a trustworthy tool for assisting medical professionals in data-driven clinical decision-making.

N. J, Deekshitha U, K. V · 0 citations
Open access Aug 2026

A Reliability-Aware and Interpretable Machine Learning Framework for Diabetes Prediction Using Structured Clinical Data

Diabetes prediction plays an important role in re-ducing long-term health risks by enabling early medical interven-tion. Although machine learning models have been widely applied to this task, many existing studies emphasise predictive accuracy while giving comparatively little attention to the reliability, interpretability, and stability of the resulting decisions. This paper develops a reliability-aware and interpretable machine learning framework for diabetes prediction from structured clinical data. Three complementary models—Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost)—are trained on the Pima Indians Diabetes dataset so that both simple linear and complex non-linear relationships are captured. Beyond conventional discrimination metrics, the reliability of the predicted probabilities is quantified using the Brier score and reliability (calibration) diagrams. Interpretability is addressed with SHapley Additive exPlanations (SHAP) at both the global (cohort) and local (individual patient) levels. Because different models frequently emphasise different predictors, we formalise a Feature Consistency Index (FCI) that quantifies the cross-model agreement of SHAP-derived feature importance and combines it with normalised importance into a single ranking score. Finally, a perturbation-based robustness analysis measures the sensitivity of each model’s output to small changes in the input record. Experi-mentally, XGBoost achieves the highest discrimination (accuracy 0.7597, ROC-AUC 0.8374), whereas Random Forest attains the best-calibrated probabilities (Brier score 0.1646), demonstrating that discrimination and reliability are not interchangeable. The FCI identifies Glucose and BMI as simultaneously the most influential and the most consistently attributed predictors, while Blood Pressure and Skin Thickness are both weak and unstable. Under a 5% Gaussian perturbation of a representative patient record, the linear and bagged models shift by less than 0.01 in predicted probability, whereas the boosted model shifts by 0.0386, revealing an accuracy–stability trade-off that a purely accuracy-driven evaluation would not expose

R. V., S. Sasirekha · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.