An agentic AI framework connecting language models to electronic health records and a biomedical knowledge graph for real-world evidence
Abstract
Accessing large-scale clinical and biomedical databases remains a significant barrier for clinicians and researchers, requiring substantial computational expertise. Agentic artificial intelligence frameworks, in which large language models (LLMs) orchestrate multi-step reasoning and query execution under interactive human supervision, offer the potential to democratize data access and accelerate evidence generation. We applied the Model Context Protocol (MCP) to integrate, within a single agentic workflow, an Observational Medical Outcomes Partnership (OMOP)-standardized electronic health record (EHR) database from an academic health system (>7 million subjects), with the Scalable Precision Medicine Open Knowledge Engine (SPOKE), a curated knowledge graph integrating relationships biomedical concepts from expert-maintained resources. Using this implementation (MedCP), we evaluated the approach across 100 benchmarking clinical research tasks, 617 biomedical factual accuracy questions (BiomixQA), an integrative case study linking clinical co-occurrence with molecular similarity, and a standardized protocol for generating and replicating real-world studies. Across the 100-task benchmark, knowledge-graph access improved mean scores overall for GPT-5.5 and Claude Opus 4.8, though the pattern differed between models; on BiomixQA, SPOKE grounding raised multiple-choice accuracy for both models, without changing true/false accuracy. In the case study, disease co-occurrence patterns extracted from the EHR correlated with molecular network similarity, surfacing mechanistic hypotheses from real-world data. Applied to the replication of published observational studies, the research protocol compressed timelines from months to hours, lowering the technical barrier to query generation and execution while study design and interpretation remained under expert supervision. An agentic AI infrastructure that combines institutional EHR data with curated biomedical knowledge via MCP can serve as a transparent, domain-grounded, supervised research assistant for real-world evidence, supporting both hypothesis generation and testing.