Skip to content
Review Open access

Development of a Framework for Deidentified Japanese Electronic Health Record Narratives Using BERT: Balancing Privacy Protection and Reproducible Entity Extraction in Real-World Data.

Apr 2026 · JMIR Medical Informatics · 0 citations · 9 references
Medicine

TL;DR

This framework enables the generation of deidentified Japanese EHR narratives that preserve contextual structure for auditing while supporting structured entity extraction, thereby addressing the trade-off between privacy protection and reproducibility in real-world data research.

Abstract

Background

Pharmacoepidemiologic studies using real-world data often lack clinical context, much of which is embedded in free-text electronic health records (EHRs). However, sharing EHR narratives is restricted by privacy requirements, and conventional deidentification approaches may reduce analytic utility and limit auditability.

Objective

To develop and evaluate a framework that (1) extracts structured cancer-related entities from Japanese EHR free text using BERT (Bidirectional Encoder Representations from Transformers)-based natural language processing (NLP) and (2) generates deidentified analytic text that preserves contextual information to support post-hoc auditing under privacy constraints.

Methods

We conducted a retrospective observational study using the DATuM IDEA database, which integrates unstructured EHR narratives with structured records (claims, prescriptions, and procedure/surgery records) from two hospitals in Japan from January 1, 2019, through June 30, 2025. Deterministic linkage was performed within ICI (Integrated Clinical Care Informatics, Inc.) prior to deidentification. The source text population comprised all linked progress notes from eligible oncology patients. Downstream recoverability analyses were conducted within a governance-constrained analyzable text cohort derived after deidentification and preprocessing. We quantified deidentification using masking rates and residual visible character density and evaluated downstream recoverability of TNM/stage-related mentions, internal logical consistency (M1 vs Stage IV), and proxy-based concordance using structured-data treatment proxies (ATC code L for systemic anticancer therapy and procedure-name keyword searches for cancer-directed interventions). A stratified manual audit of 400 deidentified narratives was conducted primarily to inspect for apparent residual direct identifiers, with retention of intended TNM and stage outputs reviewed as a secondary sanity check.

Results

Among 51,876 patients with recorded diagnoses, 10,214 had neoplasms (ICD-10 C00-D48), and linked progress notes were available for 4,383 patients (82,863 records). The downstream analyzable text cohort (Layer C), used for downstream recoverability and proxy-based concordance analyses, comprised 3,689 patients and 38,841 documents. Of these, 1,339 patients (36.3%) and 11,914 documents (30.7%) had at least one recoverable TNM or stage mention (Layer D). Across all Layer C documents, marginal document-level recoverability was 24.2% for T, N, and M elements and 23.6% for stage. Conditional on Layer D, the corresponding recoverability rates were 0.788 for T, 0.787 for N, 0.788 for M, and 0.768 for stage. Median masking rates ranged from 9.9% to 15.0% across major cancer categories. M1 and Stage IV mention indicators showed an overall agreement of 94.7%, a positive percent agreement (PPA) of 54.8%, and a negative percent agreement (NPA) of approximately 99.0%. Proxy-based concordance analysis showed a positive predictive value (PPV) of 0.700 for Stage IV using the systemic therapy proxy and a PPV of 0.031 for early-stage classification using the procedure proxy. Manual audit demonstrated retention of intended TNM/stage-related outputs after normalization and identified no apparent residual direct identifiers in the audited deidentified narratives.

Conclusions

This framework enables the generation of deidentified Japanese EHR narratives that preserve contextual structure for auditing while supporting structured entity extraction, thereby addressing the trade-off between privacy protection and reproducibility in real-world data research. CLINICALTRIAL

Read PDF

Similar papers

Open access Sep 2026

A Scalable Method for Validated Data Extraction from Electronic Health Records with Large Language Models

Health care organizations increasingly require structured, patient-level clinical variables for treatment decisions, operational workflows, quality measurement, and clinical trial screening. Relevant information is often fragmented across heterogeneous electronic health record (EHR) systems, unstructured formats, a...

Timothy J. Stuhlmiller, AJ Rabe, Jeff Rapp et al. · 0 citations
Open access Sep 2026

Three layers and two revelations: A multidisciplinary framework for curating real-world data from electronic patient records reveals a decade of breast screening performance.

OBJECTIVES To develop and validate a scalable, semi-automated framework for extracting high-granularity research data from legacy Electronic Patient Records (EPR), using a decade of family history breast screening as the exemplar. METHODS Our multidisciplinary team developed a three-layer architecture distinguishing...

Richard Sidebottom, Donna L. Webb, Des Cambell et al. · 0 citations
#artificial intelligence Review Aug 2026

Review Before Trust: Source-Grounded Integrity Gates for AI-Assisted Personal Health Records

An evidence-gated trust-promotion model that keeps generated data provisional until a deterministic monitor verifies it against the source document, demonstrating the technical feasibility of an enforceable boundary that prevents generated claims from authorizing their own reuse in a longitudinal health record.

Nora Girda, Adrian Groza · 0 citations
Review Open access Oct 2026

Guideline-grounded large language models for extracting genome-informed clinical recommendations from electronic health records.

BACKGROUND Precision medicine requires delivery and tracking of genome-informed risk assessments (GIRAs), often documented in unstructured electronic health record (EHR) notes, making large-scale evaluation reliant on labor-intensive manual chart review. Large language models (LLMs) offer a promising approach to automa...

Yi Xin, K. W. Davis, Wu-Chen Su et al. · 0 citations
Review Open access Aug 2026

A Human-Governed Clinical Informatics Framework for Safe AI-Assisted Mental Health Counseling: Secondary Framework Development and Requirement Mapping Study

AI-assisted mental health counseling should be implemented as a governed clinical information workflow rather than as an autonomous diagnostic or documentation pathway, to establish clinical safety or clinical effectiveness.

Mi-Ae Yang, Kang-Su Ha · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.