Skip to content
Review Open access

Clinical Text De-Identification Beyond PHI Detection: A Scoping Review

Jul 2026 · ET Journal · Vol 9, pp. 22 · 0 citations · 57 references

TL;DR

Clinical text de-identification cannot be reduced to named-entity recognition and local hybrid or transformer pipelines remain the most defensible baseline for routine PHI detection; LLMs are better suited to defined tasks in augmentation, transformation, and assurance.

Abstract

Background: Removing explicit protected health information (PHI) does not necessarily make a clinical narrative non-identifiable. Rare diagnoses, social circumstances, locations, temporal patterns, and treatment trajectories may still permit re-identification. Methods: This PRISMA-ScR-informed scoping review combined PubMed/MEDLINE, OpenAlex, and targeted NLP searches with OpenAlex title-and-abstract, Web of Science, and Scopus sensitivity checks. Eligible publications addressed clinical or biomedical text de-identification from 1 January 2018 to 31 March 2026. Results: The searches yielded 275 records and 258 unique records after deduplication. Eligibility assessment produced 110 core records, 11 background or review records, and 22 exclusions. All 143 eligibility-stage records underwent independent double screening and reconciliation; structured extraction covered all 110 core records. Abstract screening of 47 additional Scopus candidates found no task function or implementation paradigm outside the proposed classification. The resulting two-axis framework separates a method’s role in the workflow from its technical implementation. Conclusions: Clinical text de-identification cannot be reduced to named-entity recognition. Local hybrid or transformer pipelines remain the most defensible baseline for routine PHI detection. LLMs are better suited to defined tasks in augmentation, transformation, and assurance, with local validation and explicit control of data exposure.

Read PDF

Similar papers

Review Open access Apr 2026

Development of a Framework for Deidentified Japanese Electronic Health Record Narratives Using BERT: Balancing Privacy Protection and Reproducible Entity Extraction in Real-World Data.

This framework enables the generation of deidentified Japanese EHR narratives that preserve contextual structure for auditing while supporting structured entity extraction, thereby addressing the trade-off between privacy protection and reproducibility in real-world data research.

Nobuo Mochizuki, Atsushi Ikeda, Hideyuki Ohshima et al. · 0 citations
Open access Sep 2026

LLM-enabled Natural History Study Analysis to Support Rare Disease Research

Background: Rare diseases affect an estimated 300 million people worldwide, yet the research needed to guide diagnosis and treatment is often fragmented across multiple unstructured literature sources. Natural history studies (NHS) are a key source of this evidence, but manually extracting structured information from N...

Kevin Li, E. Sid, Qian Zhu · 0 citations
Open access Sep 2026

Resource Overlap and Reporting of Independence in Public Prostate Cancer Research

Importance Public prostate cancer data may appear under different repository identifiers or releases while representing overlapping patients, specimens, or sample records. The implications for reported analytic independence require assessment at both resource and publication levels. Objective To characterize supported...

W. Yue, A. Tewari, B. Padanilam · 0 citations
Review Open access Sep 2026

Reproducible study identification and selection in systematic reviews: barriers and paths forward

Abstract Systematic reviews are foundational to evidence-based practice, but the reproducibility and auditability of study identification and selection remain persistent challenges. Existing reporting standards, such as Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 and PRISMA-Search e...

Li-Feng Lin, Xing Xing, Chong Wu et al. · 0 citations
Review Open access Sep 2026

Dynamics of Medical Terminology Documentation Quality in the SNOMED CT Era: A Literature Review

Introduction: Accurate medical terminology recording is a cornerstone of medical record quality and health data validity, as mandated by Indonesian Ministry of Health Regulation No. 24 of 2022, which requires diagnosis coding to use the latest ICD-10 classification. Various studies in Indonesia still report low levels...

Anisa Munifatun Nabillah, Untung Slamet Suhariyono, A. Rusdi · 0 citations
Review Open access Sep 2026

Automated Extraction of Genetic Eligibility Criteria from Clinical Trial Records Using LLMs - A Technical Case Report.

INTRODUCTION Accurate interpretation of clinical trial eligibility criteria is essential for applications such as patient-trial matching and clinical decision support, particularly in precision oncology. However, relevant information, including genetic mutation requirements, is typically embedded in unstructured text w...

Georg Mathes, S. Berger, Stefan Sigle · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.