Clinical Code Mapping with LLM Tool Use: A Pilot for Automated Data Extraction of Medication and Diagnosis Information from Unstructured Clinical Notes.
Sep 2026· Studies in Health Technology and Informatics· Vol 340, pp.
172-178
· 0 citations
Medicine
TL;DR
LLMs are suitable for information extraction of medications from clinical notes for use in research databases, however, for a clinical setting where the treatment of patients would be dependent on LLM performance, the current state-of-the-art open weight models are not accurate enough.
Abstract
INTRODUCTION
Accurate clinical coding is fundamental to large-scale epidemiological studies, hospital billing, and the development of robust clinical decision support systems. Conventional methods for structured data extraction often rely on manual curation, which is prohibitively labor-intensive. Goal of this project is to determine whether current state-of-the-art open-weight LLM models are suitable for extraction of structured data from non-English (German) clinical notes.
Methods
We anonymized 35 German doctor's notes of five patients from our hospital and developed one pipeline to extract and map medications and two for diagnoses. The latter compares a RAG based approach with an agentic AI. We ran these using three open-weight LLMs on a local GPU-PC.
Results
The F1 scores for diagnoses do not exceed 0.12. If we instead consider mapping to the broad category, then the F1 score increases to 0.18. For medications, the F1 score is as high as 0.78 and even 0.95 if we consider trivial name extraction only.
Discussion
For trivial name extraction of medications, every encountered mistake is explainable. Due to limitations in the nature of the task, it is infeasible to expect a perfect score of 1 in any of the coding scenarios. Further problems in LLM output and parsing are addressed.
Conclusion
LLMs excel at extraction. They are suitable for information extraction of medications from clinical notes for use in research databases. However, for a clinical setting where the treatment of patients would be dependent on LLM performance, the current state-of-the-art open weight models are not accurate enough.
Abstract Objective Clinical narrative provides a unique window into provider reasoning and attribution for automated diagnosis assignment, but large language models (LLMs) have traditionally not performed well at medical coding. We evaluate a reproducible method for automated diagnosis assignment using LLMs in clinical...
H. Razzaghi, Nhat Nguyen, M. Pargi et al.· JAMIA Open· 0 citations
Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.
L. Barrett, N. Joshi, A. S. North et al.· medRxiv· 0 citations
Clinical information required for surgical data science (SDS) is frequently embedded in unstructured text. We developed and evaluated a reproducible pipeline for selecting locally deployed open-weight large language models (LLMs) for binary symptom annotation. In this retrospective single-center study, 1,100 German eme...
Jonas Henn, Alisa Stoll, P. Feodorovici et al.· Scientific Reports· 0 citations
Since the introduction of the Patient Rights Act, patients in Germany have gained legal access to their medical records, including clinical notes. However, these documents are typically written for healthcare professionals and are often difficult for patients to understand due to specialized terminology, abbreviations,...
M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al.· Studies in Health Technology...· 0 citations
Manual clinical DNA variant classification is the bottleneck of every clinical and research rare disease workflow. The process typically requires a curator to assemble evidence from numerous databases, weigh 28 criteria, reconcile competing evidence, and produce a defensible case for the final classification. Additiona...
Jamie-Lee M. Thompson, Debjani Das, Sally L. Dunwoodie et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.