The electronic MEdical Records and GEnomics (eMERGE) network successfully generated and returned comprehensive risk profiles using logic and data specific to 11 conditions in a secure and semi-automated fashion employing a customized REDCap database.
Abstract
Abstract Objective To describe the development and implementation of an automated platform for genomic risk prediction that integrates multiple data types. Materials and methods Using the REDCap infrastructure, we constructed the R4 (Recruitment, Results, and Risk Reduction) platform to intake data from clinical sites, partner laboratories, participant surveys, and electronic health record (EHR) data across 13 institutions. Results The R4 Portal successfully integrated data to generate genome-informed risk assessments (GIRAs) across 11 conditions for a 23 840 person cohort. Testing phases and quality control led to network-wide protocols ensuring consistency and accuracy. Discussion As the science of estimating disease risk evolves, standardized and high-throughput methods of collecting and manipulating complex data are required. Platforms should be open-source, modular, and reusable, ensuring flexibility, security, and integration across healthcare environments. Conclusion The electronic MEdical Records and GEnomics (eMERGE) network successfully generated and returned comprehensive risk profiles using logic and data specific to 11 conditions in a secure and semi-automated fashion employing a customized REDCap database.
Background Inborn errors of immunity are rare, genetically heterogeneous disorders requiring coordinated clinical, laboratory, and genetic evaluation over time. Data are often fragmented across records, laboratory systems, and genomic reports, limiting longitudinal analysis and coordinated care, particularly in the Middle East and North Africa, where structured rare disease data infrastructures remain limited. Objective To develop a Research Electronic Data Capture–based data management framework for inborn errors of immunity and demonstrate its use in a prospective multi-site setting. Methods A Research Electronic Data Capture–based framework was developed at the College of Medicine and Health Sciences, United Arab Emirates University. Modular instruments captured consent, demographics, biospecimen processing, laboratory workflows, and genetic findings within a longitudinal structure. Data dictionaries, validation rules, and conditional logic ensured data quality. The framework was deployed across participating sites for prospective data collection. Results The framework enabled integrated longitudinal documentation of enrollment, biospecimens, and genetic testing. It was implemented across two clinical sites and used to enroll patients with suspected or confirmed inborn errors of immunity. The platform supported standardized cross-site data capture and monitoring of genetic findings, including automated flagging of variants of uncertain significance. Conclusion This study demonstrates the development and early multi-site implementation of a Research Electronic Data Capture–based framework for inborn errors of immunity. By enabling standardized integration of clinical, laboratory, and genetic data, the platform supports data quality, cross-site collaboration, and tracking of evolving diagnoses. It provides a scalable foundation for rare disease research and may support improved clinical decision-making.
M. Ahmed, A. A. Bousfiha, F. Almarzooqi· Frontiers in Immunology· 0 citations
This Review highlights AI and ML frameworks for integrating genomic, multi-omics, and EHR data, and discusses how these approaches are reshaping genomics research as well as clinical practice.
Rasika Venkatesh, M. Ritchie· Nature reviews genetics· 0 citations
Background Phenotype-genotype associations underpin precision medicine by enabling disease prevention, early diagnosis, risk stratification, therapeutic target discovery, and personalized treatment. However, the rapid growth of scientific evidence has made manual curation of these associations increasingly labor-intensive, time-consuming, and incomplete. Large Language Models (LLMs) offer a potential path to scalable genomic generation and synthesis of this knowledge, but their ability to accurately identify phenotype-genotype associations and the extent to which these outputs are supported by established genomic knowledge bases remain unclear. Materials and Methods Four LLMs, Claude Sonnet 4.6, DeepSeek V4 Flash, Gemini 3 Flash Preview, and GPT-5.5, were benchmarked on six zero-shot task categories covering forward and reverse phenotype-gene and phenotype-SNP generation. A total of 4,196 associations were identified from curated inputs and evaluated through a multistage external verification pipeline comprising phenotype normalization, ontology mapping, genomic identifier validation against Ensembl, and evidence verification using both the GWAS Catalog and OMIM. Associations were assigned a fused evidence level of strong, moderate, weak, or none. Results Overall, 74.19% of generated associations were matched to at least one external genomic knowledge base; 9.15% received strong support and 54.46% moderate support. Phenotype-gene associations were more verifiable than phenotype-SNP associations (strong or moderate: 67.19% vs 54.06%). Among existing associations, Claude Sonnet 4.6 achieved the highest overall strong or moderate rate (69.2%), followed by GPT-5.5 (65.1%), DeepSeek V4 Flash (61.7%), and Gemini 3 Flash Preview (56.9%). Conclusion LLMs can support scalable generation of candidate phenotype-genotype associations. Performance varied substantially by relation type and was lower for SNP-level and rare disease associations, highlighting both the limitations of current genomic resources and the need for rigorous validation pipelines.
Caiwan Sun, Yi Xin, Sarah Zeng et al.· bioRxiv· 0 citations
One of the difficulties of the widespread application of genomic medicine is the sharing of patient genomic data quickly and securely with other health care systems. This review explores current methods, technologies and governance structures for the sharing of genomic data, as well as national genome initiatives and the genomics industry to offer a broad overview. The aim is to determine the existing practice, to discuss important technical and regulatory issues and to present practical suggestions that can facilitate future development. The review is carried out using the methodology for scoping reviews developed by Arksey and O'Malley, and information from peer-reviewed literature, gray literature, national genomics programs and industry reports. Four research questions were posed: How does the sharing of genomic data occur in clinical practice, in research settings, in national genomic programs and in commercial genomics companies? 37 relevant studies were identified and analyzed, and seven key components for effective genomic data-sharing systems were identified. These studies included clinical implementation models, research models, as well as considering ethical, legal or policy issues. Given the results, there are several technologies and frameworks to share clinical genomic data, however, there is not as wide-spread a large scale implementation as desired. Scalable infrastructures, inequities in clinical/academic genomic systems and different degrees of genomic medicine uptake in health services are key challenges. The review proposes four recommendations for future directions to improve the consistency, security and interoperability of data sharing from the genome and to promote greater uptake of the use of the clinical genome and innovation in precision medicine.
Unknown authors· ITM Web of Conferences· 0 citations
Background/Objectives: Obtaining timely access to detailed clinical trial data is not always straightforward. Privacy requirements, governance processes, and study-specific eCRF configurations can delay access, particularly during study start-up, when teams need realistic data to develop and test validation rules, reporting pipelines, and centralized monitoring tools. Methods: We developed SYNDATA, a modular framework that generates synthetic clinical trial datasets conforming to a target electronic case report form (eCRF). The framework combines study metadata from the Medidata Rave Architect Loader Spreadsheet (ALS) with selected empirical patterns learned from a closely matched reference study. It constructs patient-specific timelines from the ALS visit matrix, generates module-specific records using Bayesian networks for selected categorical dependencies and density-based methods for numeric and temporal variables, and applies postprocessing for counters, dictionary coding, and conditional missingness. A configurable Noise Tool injects controlled and reproducible data defects, including timeline inconsistencies, visit-window violations, numeric threshold exceedances, randomization or eligibility conflicts, and structural collisions, to stress-test downstream validation logic. Results: Synthetic and source data were compared descriptively using Jensen–Shannon distance, Cramér’s V, and representative plots across selected domains. Several binary operational fields showed close descriptive agreement, whereas agreement was weaker for some more complex, multi-category safety- and medication-related variables. The evaluation covered a representative subset and does not establish uniform fidelity, formal statistical equivalence, or clinical validity across all generated domains. Conclusions: SYNDATA supports reproducible generation of ALS/eCRF-conformant datasets for operational validation, reporting development, and workflow testing before real trial data are available. Its demonstrated value is limited to the evaluated operational use case; fidelity is variable-specific, and more complex domains require further development and validation.
Unknown authors· Healthcare· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.