Jul 2026· Journal of Vascular Diseases· 0 citations· 46 references
TL;DR
A Large Language Model-guided retrieval-aware framework for generating realistic synthetic EHR data on a large scale to train ML models to predict cardiovascular risks accurately and achieves high performance in predicting CVD.
Abstract
Background: Increasing access to Electronic Health Records (EHRs) has enabled the development of Machine Learning (ML) models to predict early cardiovascular disease (CVD) risk. Nevertheless, real EHR data is usually limited in availability and distribution because of confidentiality issues, legal limitations, and insufficient access. Synthetic data generation has become a promising approach to overcome such difficulties. Objectives: This study proposes a Large Language Model (LLM)-guided retrieval-aware framework for generating realistic synthetic EHR data on a large scale to train ML models to predict cardiovascular risks accurately. Methods: The framework uses a Tabular Denoising Diffusion Probabilistic Model to learn the underlying distribution of the original dataset and generate an initial synthetic dataset. To improve the clinical plausibility of the generated data, a knowledge-guided refinement module with LLaMA 2 13B combined with Retrieval-Augmented Generation (RAG) is introduced. The LLM analyzes statistical trends in real and synthetic data while retrieving relevant medical information to detect and correct clinically implausible correlations. Results: Experimental findings show that the optimized synthetic dataset preserves important statistical features of the original data and achieves high performance in predicting CVD. Conclusions: Thus, the framework offers a scalable, privacy-conservative method for generating realistic synthetic healthcare data suitable for medical research.
This study aims to design and evaluate a Model Context Protocol–enabled Custom GPT framework that integrates ML-based predictive models with external healthcare systems to deliver context-aware, explainable, and clinically actionable heart disease risk predictions and paves the way for broader adoption of agentic AI in clinical workflows.
Neha Gupta, Bhawna Singla· Journal of Electronic &...· 0 citations
Diabetes is a major risk factor for the development of cardiovascular issues which contribute to cardiovascular disease (CVD) being a leading cause of mortality worldwide. However, traditional machine learning methods are not widely adopted in healthcare systems because they lack interpretability, which is important for early and accurate CVD risk prediction and for ruling out effective clinical intervention. In this research, a hybrid architecture is proposed that incorporates diabetes related datasets as well as explainable artificial intelligence (XAI) methodologies that could improve the prediction power and transparency of the models. The proposed approach combines different datasets at the level of features and includes rigorous data pre-processing to detect metabolic and cardiovascular risk factors. Some of the significant clinical parameters are age, BMI, glucose, cholesterol, and blood pressure. These are standardized to create a single dataset which may be utilized for predictive modelling. The employment of two XAI approaches, SHAP (SHapley Additive Explanations) with tree-based ensemble models and integrated gradients with transformer based topologies, ensures both performance and interpretability. The technique improves confidence and usefulness in clinical settings by offering accurate predictions and explanations for the model’s judgments that are relevant to the circumstance. It is also utilized for visual investigation of clinical correlations of diabetes and cardiovascular disease and identify crucial risk variables. The results suggest that merging explainability approaches with powerful machine learning can considerably boost early identification and risk assessment. The proposed approach contributes to enhanced healthcare decision-making, offering a scalable, interpretable and dependable solution for cardiovascular disease prediction.
K. Deepthi, P. Bhargavi· International journal of com...· 0 citations
This Review highlights AI and ML frameworks for integrating genomic, multi-omics, and EHR data, and discusses how these approaches are reshaping genomics research as well as clinical practice.
Rasika Venkatesh, M. Ritchie· Nature reviews genetics· 0 citations
The proposed GPT2-based table-to-text framework provides a practical and clinically interpretable approach for disease prediction from limited structured healthcare data and demonstrates strong potential for early risk detection, transparent clinical decision support, and reliable deployment in real-world low-resource healthcare environments.
S. Bin Akter, S. Akter, D. Eisenberg et al.· medRxiv· 0 citations
SYNRARE is a graphical user interface based on the Synthea framework that enables easier modification and generation of synthetic Electronic Health Records of RD patients, which differ only to a definable degree from patients with common diseases, thereby enabling the benchmarking and testing of algorithms under controlled technical conditions.
Nicolai Dinh Khang Truong, Richard Rottger· arXiv.org· 0 citations
This approach combines semantic understanding of clinical narratives with structural modeling of patient-disease-treatment relationships and successfully validates synthetic EHR data utility for privacy-preserving healthcare AI development while addressing critical requirements necessary for clinical decision support system.
U. Luke, P. Asuquo, Victor Anaga et al.· E3S Web of Conferences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.