Skip to content
Review Open access

AI-Assisted Clinical Data Abstraction From Electronic Health Records: Retrospective Concordance Study

Jul 2026 · JMIR Formative Research · Vol 10, pp. e96755-e96755 · 0 citations · 25 references
Medicine

TL;DR

The findings support the feasibility of AI-assisted abstraction workflows, although further validation across larger and more diverse datasets is needed.

Abstract

Abstract Background Manual chart abstraction from electronic health records is a critical step in clinical outcomes research but is time-intensive and prone to human error. Advances in artificial intelligence (AI), particularly large language models, offer the potential to automate the extraction of structured data from unstructured clinical documentation with improved efficiency and consistency. Objective This study aimed to evaluate the accuracy and efficiency of an AI-assisted approach for extracting patient-reported outcomes from clinical notes compared with traditional human abstraction. Methods We conducted a retrospective study of 26 patients treated with low-dose radiation therapy for osteoarthritis. Human reviewers abstracted numeric rating scale (NRS; 0‐10) pain scores at baseline, the end of treatment, and the first follow-up, and von Pannewitz score (VPS; 0‐4) improvement scores at posttreatment time points. A HIPAA (Health Insurance Portability and Accountability Act)–compliant generative pretrained transformer–based AI system was prompted to extract the same end points from clinical notes. Concordance was assessed using exact match rates, the intraclass correlation coefficient for the NRS, and weighted Cohen κ for the VPS. The time required for AI vs manual abstraction was recorded. The AI system was not trained or fine-tuned on study data, and performance was evaluated directly against human abstraction to reflect real-world deployment. Results The AI system demonstrated high concordance with human abstraction, achieving an exact match rate of 92% for the NRS (95% CI 84‐96; intraclass correlation coefficient=0.96) and 94% for the VPS (95% CI 84‐98; κ=0.91). All discrepancies were minor, and no spurious values were generated. The AI system identified 1 clinically relevant data point missed during manual review. Average abstraction time per patient decreased from approximately 30 minutes to 2 minutes, representing time savings of >90%. The system also captured trends in analgesic use, but these results were not statistically significant, including reductions without escalation. Conclusions AI-assisted data abstraction demonstrated high concordance with human review in this single-institution cohort while substantially reducing the time requirements. These findings support the feasibility of AI-assisted abstraction workflows, although further validation across larger and more diverse datasets is needed.

Read PDF

Similar papers

Open access Mar 2026

Generating Patient Documents from Electronic Health Records Using Generative Artificial Intelligence: A Feasibility Study in a Japanese Cancer Center

Abstract Objectives Clinical documentation consumes substantial clinician time, potentially detracting from patient care. Generative artificial intelligence (AI) may support drafting discharge summaries and patient referral documents, but feasibility in non-Western-language oncology settings using real-world electronic health record (EHR) data remains insufficiently evaluated. This study assessed feasibility in a Japanese cancer hospital using an enterprise AI system. Methods Medical records from 61 consenting adult patients at Chiba Cancer Center were analyzed. Although the plan aimed at comprehensive EHR data, actual input was limited to extractable text (physician notes, nursing records); structured laboratory data and imaging, endoscopy, and pathology reports were not directly used, and existing summaries and external referrals were excluded to avoid information leakage. Data were converted to JavaScript Object Notation; GaiXer generated 31 discharge summaries and 30 referral documents. Four evaluators scored them; ≥80/100 was an exploratory threshold for draft-level practical utility. Feedback drove one refinement cycle. Results Generated documents scored approximately 60 to 70. A score ≥80 was reached by 9 of 31 discharge summaries in each evaluation; for referrals, none reached the threshold initially, whereas 5 of 30 did after refinement. Discharge summary scores did not substantially improve; referral scores did. Raw percent agreement among three nonphysician evaluators was high, although chance-corrected agreement varied. Wilcoxon signed-rank tests showed no significant change for discharge summaries ( p  = 0.866) but significant improvement for referrals ( p  = 0.006). Conclusion This feasibility study suggests AI may support drafting these documents in a secure environment using real-world Japanese EHR data, although the generated documents did not consistently reach the predefined threshold for draft-level utility. Findings should not be interpreted as demonstrating workload reduction or maximum performance under ideal data conditions. Future studies should evaluate larger datasets, multiple institutions and models, blinded evaluations, actual editing time, clinician acceptance, and workflow impact.

N. Michihata, Hiroshi Ishii, H. Tsujimura et al. · 0 citations
Open access Jul 2026

Expert-Rated Documentary Quality of AI-Assisted Hospital Discharge Reports: A Retrospective Paired Comparison with Physician-Written Reports

AI-assisted discharge reports received higher expert-rated documentary quality scores in a non-blinded paired evaluation across most evaluated dimensions, supporting the need for a supervised hybrid model in which AI generates the initial draft while the clinician mandatorily validates sensitive content.

Daniela Velásquez-Villegas, Toni Alonso Solís, Alex Trejo-Omeñaca et al. · 0 citations
Review Aug 2026

Clinical evaluation of a vision-language model for optimizing triage and clinical workflows in critical care.

The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and patient monitoring methods.

I. Strechen, P. Krishnan, O. Kilickaya et al. · 0 citations
Open access Jul 2026

Neuro-Symbolic AI for Automated Pathology Quality Measurement

Background. Clinical quality measurement often relies on manual abstraction of medical records, an approach that is costly, burdensome, and often infeasible for measures requiring interpretation of narrative text; these constraints have shaped measure development itself, filtering out clinically important measures that are too difficult to operationalize. We evaluated whether neuro-symbolic artificial intelligence (NSAI), which combines large language model extraction with symbolic reasoning, could reliably abstract complex quality measures from narrative pathology reports. Methods. The NSAI system decomposes each measure into atomic questions and is aligned to real-world reports through case-based refinement, an iterative human-in-the-loop process. Using 2,000 independently double-abstracted reports, we compared NSAI-based abstraction against trained human abstractors across four pathology quality measures established by the College of American Pathologists. Results. The NSAI system's agreement with the adjudicated gold standard (Cohen's kappa = 0.95) matched or modestly exceeded that of the trained human abstractors measured against the same standard (kappa = 0.92), with particularly strong performance on Gastrointestinal Metaplasia (CAP 43). In component analyses, case-based refinement drove the largest accuracy gains (up to 0.25 in kappa), whereas architectural decomposition primarily reduced performance variance across language-model backends more than tenfold, a property essential for clinical deployment. Conclusions. These findings suggest that automated abstraction could enable census-level quality measurement, reduce reporting burden, and expand the range of clinically meaningful measures that can be operationalized from narrative clinical documentation.

F. Brann, L. Tadele, C. Skau et al. · 0 citations
Review Jul 2026

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records

Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world discharge summaries, and to identify recurring failure modes that limit reliability at scale. Materials and Methods: We applied a two-stage LLM pipeline---open-ended candidate identification (Gemini 2.5 Pro) followed by context-grounded verification (Gemini 2.5 Flash)---to 3,000 randomly sampled MIMIC-IV-Note discharge summaries. A subset of the pipeline output was then reviewed manually by clinical experts. Results: Our pipeline surfaced 3,460 candidate inconsistencies, affecting 69.7% of admissions. Representative examples spanned demographics, allergies, procedures, diagnoses, laboratory, medications, and care-planning domains, with direct implications for clinical reasoning or patient safety. Expert review also revealed recurring failure modes that arise when verification requires temporal reasoning, evolving-diagnosis context, or knowledge of outpatient-prescribing conventions the model does not natively possess. Discussion: Detection is highly context-dependent: many flagged pairs require anchoring each statement to its source section and clinical domain, then assessing whether the conflict reflects a true contradiction or missing context. We propose a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsistency axis. Conclusion: This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis.

Jiankang Lu, Panyu Chen, Miriam Treggiari et al. · 0 citations
Review Open access Aug 2026

Explainable Artificial Intelligence for Tabular Data in Healthcare: A Systematic Review of Methods, Evaluation, and Applications

This systematic review provides a comprehensive analysis of XAI methods specifically applied to tabular healthcare data for classification tasks, revealing that SHAP remains the dominant post-hoc method, achieving strong model fidelity but showing inconsistent alignment with clinical expert reasoning.

Angelower Santana-Velásquez, M. B. Salazar-Sánchez · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.