Jul 2026· Journal of Biomedical Semantics· 0 citations
Medicine
TL;DR
Fine-tuning-based approaches outperform models without task-specific adaptation and prompt-based strategies, and domain-specific fine-tuning remains the most effective strategy for real-world clinical text de-identification.
Abstract
Background
Unstructured clinical narratives in electronic health records contain essential information for healthcare delivery and research. However, the presence of personally identifiable information poses significant privacy risks, which limit the secondary use of data. Therefore, reliable automated de-identification is a prerequisite for the reuse of clinical texts. This study aims to evaluate different strategies for Spanish clinical text de-identification by assessing the performance of language models on synthetic and real-world datasets.
Methods
This study presents a comparative evaluation of de-identification approaches using two datasets in Spanish: MEDDOCAN, a publicly available synthetic corpus, and ObstEHR, real-world clinical narratives from an obstetrics department. Three de-identification strategies were assessed: (i) inference with a pre-trained task-specific model used as-is, (ii) prompt-based inference using locally deployed large language models (LLMs), and (iii) fine-tuned task-specific models. Model performance was evaluated using precision, recall, and F1 score on test sets. In addition, a text preservation metric was introduced to assess prompt-based LLMs, and the impact of training set size was analyzed using progressively larger subsets of annotated ObstEHR data.
Results
Models without fine-tuning and prompt-based LLMs showed limited performance, with macro-averaged F1 scores ranging from 0.14 to 0.5 on ObstEHR and from 0.19 to 0.5 on MEDDOCAN. Fine-tuned models achieved higher performance, reaching macro-averaged F1 scores of up to 0.956. Learning curve analyses showed consistently high precision and gradual improvements in recall as the amount of training data increased, with high performance achieved with moderate amounts of annotated data.
Conclusion
Fine-tuning-based approaches outperform models without task-specific adaptation and prompt-based strategies. Despite their flexibility in generative settings, LLM-based prompt strategies show limited reliability and information preservation in clinical de-identification. Prompt-based LLMs make slight textual modifications that cause token misalignment, leading to a subsequent decrease in evaluation metrics. Therefore, domain-specific fine-tuning remains the most effective strategy for real-world clinical text de-identification.
Since the introduction of the Patient Rights Act, patients in Germany have gained legal access to their medical records, including clinical notes. However, these documents are typically written for healthcare professionals and are often difficult for patients to understand due to specialized terminology, abbreviations,...
M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al.· Studies in Health Technology...· 0 citations
The findings support the feasibility of applying LLM-based natural language processing tools in resource-limited, non-English healthcare settings and should assess emerging high-parameter models and explore additional clinical domains.
Breno Gabriel Araújo Sampaio de Jesus, Tomaz Castrillon Figueiredo, Clariele de Almeida Pereira et al.· Cadernos de Saúde Pública· 1 citation
Clinical notes contain personally identifiable information (PII), restricting reuse for research and medical AI, especially when data cannot leave an institution. We developed MedDeID, an on-premises framework combining in-house annotation and synthetic-note generation with model training, inference, pseudonymisation a...
Stig Hellemans, T. Stroobants, E. Scheurwegs et al.· 0 citations
Background: Free-text notes in electronic health records (EHRs) contain fine-grained psychiatric information that is essential for psychiatric research and clinical care, and often absent or under-recorded in structured codes alone. Clinical natural language processing (cNLP) can support extraction of this information...
X. Xue, C. Frydman-Gani, A. Arias et al.· medRxiv· 0 citations
INTRODUCTION
Clinical narratives in electronic health records frequently contain clinical expressions describing medical conditions. Their free-text format limits interoperability and automated processing. Medical concept normalization (MCN) addresses this challenge by mapping textual expressions to standardized termin...
Helena Adam, Akhila Abdulnazar, Roland Roller et al.· Studies in Health Technology...· 0 citations
Effective detection of ADEs in clinical notes may benefit from NLP models that approximate the clinical reasoning of health care providers as models evolve.
Alan Katz, Abhishek Dhankar, Gillian Fransoo et al.· Journal of Medical Internet...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.