Skip to content
Open access

Enhancing Large Language Models for Identifying and Prioritizing Important Medical Jargons From Electronic Health Record Notes Using Data Augmentation: Comparative Study

Feb 2025 · JMIR AI · Vol 5 · 1 citation · 120 references
Medicine Computer Science

TL;DR

This study evaluated both closed-source and open-source large language models for extracting and prioritizing medical jargon from EHR notes relevant to individual patients, leveraging prompting techniques, fine-tuning, and data augmentation and found that model performance could deviate largely based on prompting styles.

Abstract

Background OpenNotes allows patients to access their electronic health record (EHR) notes through online patient portals. However, EHR notes contain abundant medical jargon, which can be difficult for patients to comprehend. One way to improve comprehension is by reducing information overload and helping patients focus on the medical terms that matter most to them. Objective This study aimed to evaluate both closed-source and open-source large language models (LLMs) for extracting and prioritizing medical jargon from EHR notes relevant to individual patients, leveraging prompting techniques, fine-tuning, and data augmentation. Methods We evaluated the performance of closed-source and open-source LLMs on a dataset of 90 expert-annotated EHR notes. We tested various combinations of settings, including (1) general and structured prompts, (2) zero-shot and few-shot prompting, (3) fine-tuning, and (4) data augmentation. To enhance the extraction and prioritization capabilities of open-source models in low-resource settings, we applied data augmentation using GPT-4o and integrated a ranking technique to refine the training process. Additionally, to measure the impact of dataset size, we fine-tuned the models by incrementally increasing the size of the augmented dataset from 10 to 9995 and tested their performance. The effectiveness of the models was assessed using 10-fold cross-validation, providing a comprehensive evaluation across various settings. We report the F1-score and mean reciprocal rank for performance evaluation using two different string matching algorithms (relaxed string matching and Jaccard Index). We also conducted an error analysis classifying the erroneous outputs from the models. Results Our results show that open-source models achieved the highest performance, particularly when using fine-tuning with a gold-standard dataset. Under Jaccard Index–based string matching, DeepSeek 8B set the benchmarks with an F1-score of 0.431 (SD 0.046); similarly, BioMistral 7B showed a mean reciprocal rank of 0.577 (SD 0.109). However, under relaxed string matching, open-source models were unable to match the performance of closed-source models, even with data augmentation or fine-tuning. We analyzed our experiment from several perspectives. First, few-shot prompting did not show an advantage over zero-shot prompting in vanilla models. Second, when comparing general and structured prompts, we found that model performance could deviate largely based on prompting styles. Third, fine-tuning with a small gold-standard dataset improved performance. Finally, data augmentation yielded performance comparable to or even surpassing the fine-tuning strategy. However, it also underscored the importance of the quality of the augmented dataset. Conclusions The evaluation of both closed-source and open-source LLMs highlighted the effectiveness of prompting strategies, fine-tuning, and data augmentation in enhancing model performance in low-resource scenarios.

Read PDF

Similar papers

Aug 2026

Quality-aware multi-source data fusion and enhancement for Medical Concept Normalization using Large Language Models.

This study finds that existing MCN benchmarks present data quality issues and underexplored data fusion potential, and data quality enhancement and LLM-based controlled-variety data augmentation help alleviate overlapping phrases and long-tail issues.

Yuhan Zhou, Ruochi Li, Ana Cleveland et al. · 0 citations
Open access Aug 2026

Enhancing Low-Resource Healthcare Chatbots via Multi-Stage Data Augmentation and IndoBERT Fine-Tuning

Indonesian healthcare question-answering (QA) systems often operate in low-resource settings and must handle substantial linguistic variability in real-world user queries, including paraphrasing, informal expressions, and implicit intent. These challenges are compounded by limited annotated healthcare data and the diverse ways patients express similar medical needs in everyday language, causing QA models to rely heavily on surface-level wording. This study proposes a new multi-stage data augmentation method to improve the robustness of a healthcare chatbot based on IndoBERT Fine-Tuned. The proposed method integrates domain-specific fine-tuning with an augmentation pipeline that introduces paraphrased question variants with IndoT5, normalizes informal language, and incorporates controlled lexical variation through Part-of-Speech (POS) based verb synonym replacement while preserving medical entities, thereby expanding linguistic coverage without requiring additional manual annotation. The augmentation process preserves medical intent while generating diverse surface forms, enabling the model to learn more flexible representations of user queries. Experimental results demonstrate that the proposed multi-stage augmentation substantially improves the robustness of the IndoBERT-based healthcare question answering system. Full augmentation expands the training data to 1,424 question–answer pairs and achieves 67.37 Exact Match (EM) and 85.59 F1. Compared with training without augmentation, this corresponds to relative improvements of approximately 11.8% in EM and 7.0% in F1, indicating more reliable answer span extraction under paraphrased, informal, and lexically varied queries. These gains reflect improved alignment between conversational user input and structured healthcare information. Overall, this work highlights the importance of integrated data augmentation for enhancing low-resource Indonesian healthcare question-answering systems. By exposing the model to broader linguistic variations during training, the proposed approach supports more stable real-world performance while maintaining medical intent consistency and provides insights into remaining challenges for reliable healthcare QA deployment.

A. Rosyadi, Taufiqur Rohman, Moh. Rizki Fajar et al. · 0 citations
Open access Aug 2026

Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.

L. Barrett, N. Joshi, A. S. North et al. · 0 citations
Open access Sep 2026

Information retrieval in pre-hospital care with visualization-oriented natural-language interface via LLMs

With the popularization of Electronic Health Records (EHR), the emergency system has stored a large number of historical dispatch records, which can provide valuable insights for the optimization of current pre-hospital care. However, the inconvenient interaction manner of cur-rent information retrieval systems hinders researchers from exploring these historical records. To address this issue, we propose a novel framework that leverages the language understanding and code generation ability of Large Language Models (LLMs) to build an information retrieval system with Visualization-oriented Natural-language-based Inter-faces (V-NLI). To incorporate both domain-specific and task-related prior knowledge, we generate the instruction datasets based on the ability of closed-source LLMs in a multi-stage manner and conduct supervised fine-tuning on open-source LLMs. We also devised various mechanisms for augmenting the capabilities of open-source LLMs in query interpretation and code generation. To validate the effectiveness and generalizability of our framework, we conducted experiments on a public dataset NLV. More significantly, we performed more detailed experiments on a dataset including over 1 mil-lion pre-hospital emergency historical records in ten years. The performance of our method surpasses all baseline methods and achieves comparable results even with some SOTA closed-source models.

Xin Gao, Zheng-Ye Zhu, Xin-Yu Ma et al. · 0 citations
Review Open access Jul 2026

Tutorial: guidance on the use of large language models for medical research

This entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.

Qiao Jin, Nicholas Wan, Robert Leaman et al. · 1 citation
Open access Jul 2026

Benchmarking large language models for clinical data extraction from Portuguese medical notes in a university hospital

The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.

Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.