Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#federated learning Dataset Open access Sep 2026

Vector Database Dataset

Dataset generated from the following paper. AbstractBackground Strict privacy regulations and institutional barriers have created persistent medical data silos, limiting the development and deployment of artificial intelligence (AI) models across healthcare institutions. Although federated learning provides a potential solution, it still requires access to decentralized patient-level data and faces challenges related to interoperability, computational complexity, and institutional governance. We investigated whether a private-domain clinical vector database could enhance deep learning-based outcome prediction without exchanging raw patient data. Methods We retrospectively collected 10 years of real-world electronic health records (EHRs) from more than 40,000 cardiovascular inpatients at a tertiary hospital in Shanghai, China. The clinical cohort included patients hospitalized with cardiovascular diseases, with outcomes defined as in-hospital mortality and readmission events. Clinical text records were transformed into high-dimensional representations using a locally constructed Word2Vec-based cardiovascular clinical vector database. We developed and compared two modeling pipelines: (1) conventional pretrained language models using clinical text representations alone, and (2) hybrid models integrating pretrained language models with private-domain vector embeddings through embedding fusion. 20% of patients were independently reserved as a test cohort. Model performance was evaluated using area under the receiver operating characteristic curve (AUC), sensitivity, specificity, accuracy, calibration performance, and prediction variance. Results Pretrained language models, including ChineseBERT, RoBERTa, and MacBERT, achieved reliable baseline performance for cardiovascular outcome prediction. Incorporation of the private clinical vector database significantly improved predictive performance, with ChineseBERT- and MacBERT-based hybrid models achieving approximately 10% absolute increases in AUC, while RoBERTa demonstrated a consistent improvement of approximately 2% in AUC. The vector-enhanced models also showed improved calibration and reduced prediction variability compared with standalone language models. Conclusions This study establishes a private-domain cardiovascular clinical vector database constructed from long-term real-world EHRs and demonstrates that model–model interaction through local embedding integration can improve deep learning prediction of cardiovascular in-hospital mortality and readmission outcomes. By enabling knowledge enhancement without patient-level data sharing, this framework provides a scalable and privacy-preserving strategy to overcome healthcare data silos and facilitates the development of precision cardiovascular AI systems.

Yanan Dai · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.