This work introduces a scalable, resource-efficient, and high-performance information extraction pipeline that leverages large language models (LLMs) to address challenges of free-text clinical records and develops a multi-dimensional assessment for deployment in data extraction tasks.
Abstract
Free-text clinical records represent an untapped wealth of data for secondary use, but realising their potential is limited by resource demands necessary for accurate information extraction at scale. We introduce a scalable, resource-efficient, and high-performance information extraction pipeline that leverages large language models (LLMs) to address these challenges. Our pipeline was developed and tested using real-world dual specialist-annotated ophthalmic clinical letters, and achieved strong performance with a proprietary model in development, yielding a maximum micro-averaged F1 score of 0.954 (95% CI 0.941–0.967) for diagnosis across nine conditions through iterative prompt refinement alone, also demonstrating strong generalisability (micro-F1 0.945–0.980) in temporal validation. This approach was extended to other models in the same family and 17 LLMs from seven open-weight LLM families. Beyond performance, we develop a multi-dimensional assessment for deployment in data extraction tasks, including an error taxonomy and Pareto frontier analyses to systematically map the operational trade-offs across different LLM configurations. A robust approach to operationalisation in real-world workflows at scale may help lay the foundation for next-generation data pipelines that accelerate scientific discovery and power continuous learning health systems.
The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.
Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al.· Cadernos de Saúde Pública· 1 citation
This is the first study to benchmark SLMs on Italian EHRs and investigate the role of clinical expertise in prompt engineering, offering valuable insights for the future integration of SLMs into real-world clinical workflows.
Federica Corso, V. Peppoloni, L. Mazzeo et al.· Communications Medicine· 0 citations
VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.
Praveen Reddy, C. Mandke, Suvrankar Datta et al.· 0 citations
Abstract Background Large language model (LLM) agents capable of generating and executing statistical code from natural language may broaden access to clinical data analysis, yet which pipeline stages they perform reliably and which require expert oversight remain poorly defined. Objective This study aimed to evaluate the performance and systematic failure modes of an LLM agent across 5 stages of a clinical data analysis workflow. Methods The publicly available dataset and R script (R Foundation for Statistical Computing) were drawn from a previously published study of 12-year outcomes in 7802 patients with eyes with neovascular age-related macular degeneration at Moorfields Eye Hospital. Participants were evaluated using an LLM agent (Claude; Anthropic) across 3 interaction modes (Chat, Code, and Cowork). It was asked to perform 3 levels of data analysis practice: prompt A, to generate research questions from raw data only; prompt B, to develop a statistical analysis plan (SAP) from a high-level clinical objective, then execute it; and prompt C, to execute an analysis given an investigator-drafted SAP. Each was replicated 3 times (27 total runs). Qualitative evaluation of research question thematic coverage (prompt A), SAP completeness against a reference checklist (prompt B), and evaluation of execution outputs against validated reference values and of result text and narrative summaries against execution logs (prompts B and C) was conducted. Results The agent generated 18 clinically grounded questions spanning 7 domains; Cowork mode uniquely reached 3 thematic areas requiring data-driven methods. All 9 SAPs correctly identified the statistical framework. Kaplan-Meier estimates were near-identical across 17 completed runs. Systematic execution errors emerged: SAP quality did not predict code correctness, and within-mode errors propagated identically across independent repetitions. Result text accurately reflected execution logs in nearly all runs, though unit propagation and an undisclosed postcrash rerun were identified. Of 17 narrative summaries, 8 were fully satisfactory; 2 runs produced clinically meaningful errors. Conclusions LLM agents perform reliably for question generation and SAP drafting but require expert verification of formula composition, cohort boundary logic, and concordance computation before results are reported. Using an ophthalmology dataset as a controlled testbed, this study develops and applies an evaluation framework whose lessons are likely applicable across clinical specialties.
Unknown authors· Journal of Medical Internet...· 0 citations
Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.
L. Barrett, N. Joshi, A. S. North et al.· medRxiv· 0 citations
An open-source, human-verified workflow using large language models can accelerate electronic health record abstraction while improving accuracy and supports broader adoption of transparent artificial intelligence methods in clinical research.
Carl Jannes Neuse, Malte Janssen, S. Ibing et al.· BMC Medical Informatics and...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.