Skip to content
Review Open access

Guideline-grounded large language models for extracting genome-informed clinical recommendations from electronic health records.

Oct 2026 · JAMIA Journal of the American Medical Informatics Association · 0 citations
Medicine

Abstract

Background

Precision medicine requires delivery and tracking of genome-informed risk assessments (GIRAs), often documented in unstructured electronic health record (EHR) notes, making large-scale evaluation reliant on labor-intensive manual chart review. Large language models (LLMs) offer a promising approach to automated extraction, but their performance in genome-informed clinical contexts remains incompletely characterized.

Objectives

To evaluate guideline-grounded LLM approaches for extracting genome-informed clinical recommendations from EHR notes in the eMERGE study.

Materials And Methods

We developed an LLM-based pipeline to identify clinical recommendations associated with GIRA reports. Three LLMs (GPT-4o, LLaMA-3.3-70B, and LLaMA-3.1-8B) were evaluated using baseline prompting, guideline-aware prompting, and retrieval-augmented generation (RAG), validated against manual chart review (N = 18 participants; N = 34 documents).

Results

Baseline prompting showed limited performance (F1 = 0.50-0.54). Incorporating guideline knowledge improved F1 by 0.14-0.31 across models, with LLaMA-3.1-8B using RAG achieving the highest performance (F1 = 0.81).

Conclusion

Domain-grounded LLM approaches can support scalable extraction of genome-informed clinical recommendations from EHR data. Larger studies are needed to assess generalizability in real-world workflows.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.