Skip to content
Book Open access

Enhancing Biomedical AI Foundations: Genomic Literature Knowledge Base Boosts LLMs' Mastery of Biomedical Literature

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 9149-9158 · 0 citations · 12 references

Abstract

Peered-reviewed literature provides reliable biomedical knowledge, which is essential for large language models (LLMs) in solving complex biomedical problems. However, how literature is retrieved and presented to LLMs can significantly influence their performance on domain-specific tasks.We present the Genomic Literature Knowledge Base (GLKB) and GLKB Agent to address these challenges. GLKB is a large-scale knowledge graph containing 14.6 million relationships among 3.2 million entities from 33 million PubMed abstracts and nine curated biomedical repositories. The current release includes articles published before March 2025. It supports diverse applications, including reinforcement learning, link prediction, and semantic embedding. The GLKB Agent is an agentic architecture that seamlessly connects LLMs to GLKB. It enables autonomous retrieval, reasoning, and deep research capabilities. Our evaluations demonstrate that the GLKB Agent dramatically improves LLM performance. Eight state-of-the-art LLMs achieve average accuracy gains of 27.5% on PubMedQA-HC, 24.8% on PubMedQA-Artificial, and 6.0% on BioASQ. Ablation tests confirm the agentic architecture is particularly effective for complex reasoning tasks. During datasource ablation tests, GLKB outperforms alternative data sources including PubMed, Wikipedia, and arXiv. Beyond question-answering, GLKB agent also demonstrates deep research capabilities through test-time reasoning. It generates comprehensive reports for literature reviews and hypothesis generation. The GLKB and GLKB Agent together provide a strong foundation for next-generation biomedical AI. Access is available at: https://glkb.org.

Read PDF

Similar papers

Optimizing large language model prompts for biomedical knowledge discovery

This work presents a scalable, reproducible framework for evaluating, optimizing, and interpreting LLMs for biomedical knowledge extraction, with a focus on gene–gene regulatory relation prediction, pathway component recognition, multimodal pathway figure understanding, and automated prompt optimization.

Muhammad Azam · 0 citations
Open access Sep 2026

An agentic AI framework connecting language models to electronic health records and a biomedical knowledge graph for real-world evidence

Accessing large-scale clinical and biomedical databases remains a significant barrier for clinicians and researchers, requiring substantial computational expertise. Agentic artificial intelligence frameworks, in which large language models (LLMs) orchestrate multi-step reasoning and query execution under interactive human supervision, offer the potential to democratize data access and accelerate evidence generation. We applied the Model Context Protocol (MCP) to integrate, within a single agentic workflow, an Observational Medical Outcomes Partnership (OMOP)-standardized electronic health record (EHR) database from an academic health system (>7 million subjects), with the Scalable Precision Medicine Open Knowledge Engine (SPOKE), a curated knowledge graph integrating relationships biomedical concepts from expert-maintained resources. Using this implementation (MedCP), we evaluated the approach across 100 benchmarking clinical research tasks, 617 biomedical factual accuracy questions (BiomixQA), an integrative case study linking clinical co-occurrence with molecular similarity, and a standardized protocol for generating and replicating real-world studies. Across the 100-task benchmark, knowledge-graph access improved mean scores overall for GPT-5.5 and Claude Opus 4.8, though the pattern differed between models; on BiomixQA, SPOKE grounding raised multiple-choice accuracy for both models, without changing true/false accuracy. In the case study, disease co-occurrence patterns extracted from the EHR correlated with molecular network similarity, surfacing mechanistic hypotheses from real-world data. Applied to the replication of published observational studies, the research protocol compressed timelines from months to hours, lowering the technical barrier to query generation and execution while study design and interpretation remained under expert supervision. An agentic AI infrastructure that combines institutional EHR data with curated biomedical knowledge via MCP can serve as a transparent, domain-grounded, supervised research assistant for real-world evidence, supporting both hypothesis generation and testing.

Unknown authors · 0 citations
Preprint Aug 2026

ANCHOR-RE: An Agentic Neuro-Symbolic Framework for Grounded Biomedical Relation Extraction

Biomedical relation extraction (BioRE) extracts structured knowledge from biomedical literature for applications such as knowledge base construction and hypothesis generation. Traditional symbolic systems such as SemRep provide high precision but limited recall, while large language models (LLMs) offer stronger contextual reasoning but remain prone to false-positive predictions. We developed ANCHOR-RE, a framework that integrates ontology-guided reasoning, external knowledge grounding, and data-driven verification rules into LLM inference. We evaluated it on three BioRE benchmarks (SemRepGS, DDI, and ChemProt) using both proprietary and open-weight LLMs. To assess generalizability beyond benchmark datasets while reducing potential evaluation bias from LLM pretraining contamination, we conducted a temporal evaluation using 100 biomedical articles published in 2026. With the proprietary backbone, ANCHOR-RE outperformed direct LLM prompting, improving micro-F1 from 0.654 to 0.676 on SemRepGS, from 0.769 to 0.872 on DDI, and from 0.939 to 0.941 on ChemProt. On DDI and ChemProt, it also outperformed previously reported inference-only methods and approached fine-tuned or instruction-tuned systems without parameter updates. Similar performance gains observed with open-weight LLMs indicate that the benefits were not limited to the proprietary backbone. On the post-cutoff set, manual assessment of 500 randomly sampled predictions yielded a precision of 69%, maintaining consistent precision on previously unseen biomedical literature. Neuro-symbolic reasoning can improve the reliability of LLM-based BioRE without fine-tuning. Results across multiple benchmarks, model families, and post-cutoff literature support ANCHOR-RE as a practical training-free approach to biomedical literature mining.

Shufan Ming, Yikun Han, Gibong Hong et al. · 0 citations
#natural language process... Preprint Aug 2026

Surgical Alignment in Knowledge Graph Training for Clinical Diagnosis with Large Language Models

A systematic study spanning five KG task formulations, three training paradigms, two KGs, and three base LLMs finds that at the task level, all paradigms improve over the non-finetuned baseline, but methods with comparable in-domain accuracy show substantially different knowledge transfer behavior.

Saksham Khatwani, He Cheng, Majid Afshar et al. · 0 citations
Review 2026

Scalable Extraction and Normalization of Biomedical Knowledge from Research Literature

A knowledge graph generation pipeline that extracts subject-predicate-object triplets representing scientific claims from research articles and implements a multi-stage consolidation process focused on normalizing entities and relations, offering a robust framework for reducing ambiguity and improving the interoperability of the resulting knowledge graphs.

S. Krovvidi, Finn Vos, Laurent D. Hasson et al. · 0 citations
Dataset Open access Jul 2026

VitaGraph: building a knowledge graph for biologically relevant learning tasks

VitaGraph is presented, a comprehensive multi-purpose biological knowledge graph built by integrating and refining multiple public datasets and enabling benchmarking of graph-based models and offering the opportunity to tackle tasks such as drug repurposing, PPI prediction, and side-effect prediction, among others.

Francesco Madeddu, Lucia Testa, Gianluca De Carlo et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.