Skip to content
Review Open access

Artificial Intelligence-Generated Electronic Medical Record Summarization in Breast Surgical Oncology.

Jul 2026 · Annals of Surgical Oncology · 0 citations · 23 references
Medicine

TL;DR

Although users rated RAG-enabled GPT-4o agent-generated documentation summaries favorably on several quality domains, they frequently lacked thoroughness and occasionally contained treatment-relevant errors.

Abstract

Background

Reviewing pathology, imaging, and consultation documents in oncology can be time-consuming, particularly when records originate from external facilities in different file formats. This study aimed to evaluate the impact of a Retrieval-Augmented Generation (RAG)-enabled GPT-4o summarization agent on clinical workflows and quality of outside-record summaries in breast surgical oncology.

Methods

Initial performance evaluation of a GPT-4o/RAG agent to generate summaries of oncologic reports in 50 charts followed by a prospective pilot test of sequential cases, with each AI summary evaluated using a modified Provider Documentation Summarization Quality Instrument (PDSQI-9; 1-5 Likert scale), including dichotomized ratings (low [1-3], high [4, 5]), binomial testing, frequency and type of user-reported errors, clinician-coded error criticality (treatment-impacting vs noncritical). Pre- and post-use survey of documentation burden (NASA TLX) and user experience was performed.

Results

Among 62 cases, AI-generated summaries were rated high for accuracy, usefulness, succinctness, and source citation. Thoroughness without omission was rated low in 28 (45%) summaries. Errors were noted in 25 (40%) surveys, with 13 (52%) classified as critical (treatment-impacting). The most common error type involved imaging, reported in 17 (68%) cases. For perceived time savings, the median response was neutral, but qualitative feedback described the tool as helpful for straightforward cases and as reducing typing burden but requiring workflow adjustment and improvements for complex cases.

Conclusions

Although users rated RAG-enabled GPT-4o agent-generated documentation summaries favorably on several quality domains, they frequently lacked thoroughness and occasionally contained treatment-relevant errors. Human review and further iteration of the technology remain necessary before implementation.

Read PDF

Similar papers

Open access Mar 2026

Generating Patient Documents from Electronic Health Records Using Generative Artificial Intelligence: A Feasibility Study in a Japanese Cancer Center

Abstract Objectives Clinical documentation consumes substantial clinician time, potentially detracting from patient care. Generative artificial intelligence (AI) may support drafting discharge summaries and patient referral documents, but feasibility in non-Western-language oncology settings using real-world electronic health record (EHR) data remains insufficiently evaluated. This study assessed feasibility in a Japanese cancer hospital using an enterprise AI system. Methods Medical records from 61 consenting adult patients at Chiba Cancer Center were analyzed. Although the plan aimed at comprehensive EHR data, actual input was limited to extractable text (physician notes, nursing records); structured laboratory data and imaging, endoscopy, and pathology reports were not directly used, and existing summaries and external referrals were excluded to avoid information leakage. Data were converted to JavaScript Object Notation; GaiXer generated 31 discharge summaries and 30 referral documents. Four evaluators scored them; ≥80/100 was an exploratory threshold for draft-level practical utility. Feedback drove one refinement cycle. Results Generated documents scored approximately 60 to 70. A score ≥80 was reached by 9 of 31 discharge summaries in each evaluation; for referrals, none reached the threshold initially, whereas 5 of 30 did after refinement. Discharge summary scores did not substantially improve; referral scores did. Raw percent agreement among three nonphysician evaluators was high, although chance-corrected agreement varied. Wilcoxon signed-rank tests showed no significant change for discharge summaries ( p  = 0.866) but significant improvement for referrals ( p  = 0.006). Conclusion This feasibility study suggests AI may support drafting these documents in a secure environment using real-world Japanese EHR data, although the generated documents did not consistently reach the predefined threshold for draft-level utility. Findings should not be interpreted as demonstrating workload reduction or maximum performance under ideal data conditions. Future studies should evaluate larger datasets, multiple institutions and models, blinded evaluations, actual editing time, clinician acceptance, and workflow impact.

N. Michihata, Hiroshi Ishii, H. Tsujimura et al. · 0 citations
Review Open access Aug 2026

The daily dose: Early usability of an LLM tool for patient summaries and trial matching in radiation oncology

Background Radiation oncology workflows generate large volumes of electronic health record (EHR) data requiring daily synthesis. Large language model (LLM)-based automation is promising, but workflow-embedded implementations at scale remain limited. We describe the design and early usability and adoption evaluation of The Daily Dose (TDD), an LLM-driven system for automated clinical summarization and trial identification in radiation oncology. Materials and methods TDD delivers physician-specific email summaries each morning across three Mayo Clinic campuses using RadOnc-GPT (GPT-4o) to generate EHR-derived patient summaries and identify potentially eligible clinical trials for new or consult visits. One month post-deployment, an anonymous cross-sectional survey adapted from the System Usability Scale and Technology Acceptance Model was administered to all recipients. Results Fifty-five of 110 users responded (50%); 94.5% were in radiation oncology and 69.1% were attending physicians. Overall, 83.6% used TDD at least several times per week. Mean domain scores (5-point Likert) were 3.89 ± 1.04 for usability and satisfaction, 3.43 ± 1.24 for perceived usefulness, and 3.80 ± 1.17 for impact and future use. Satisfaction was significantly associated with perceived time savings (p < 0.001); 27% estimated saving ≥10 min daily. Internal consistency was high (α = 0.97). Free-text responses highlighted improved preparedness and patient-context awareness but noted occasional inaccuracies and imperfect trial matching. Conclusion In this early usability and adoption evaluation, a workflow-integrated LLM summarization tool was widely adopted and generally favorably perceived. These findings reflect user perceptions; objective validation of summary accuracy, trial-matching performance, and workflow efficiency is needed to establish clinical impact.

J. Holmes, F. Mastroleo, M. Borras-Osorio et al. · 0 citations
Aug 2026

Radiologically Relevant Clinical History Summarization with Large Language Models: A Multireader Performance Study.

LLMs generated radiology-relevant indications from clinical notes that were more comprehensive and factual than clinician indications, and when generated by the proprietary LLM, were ranked most useful in protocoling and imaging interpretation.

A. Serapio, Timothy L. Chen, Brian Tangsombatvisit et al. · 1 citation
Review Open access Jul 2026

A Modular Evaluation of AI-Assisted Clinical Documentation

The results support the use of this modular AI-assisted clinical documentation pipeline as a human-supervised draft-generation tool that still requires clinician review, local workflow evaluation, and prospective clinical validation before broader deployment.

Julien Delaunay, Maissaa Sarkis, Jordi Solé-Casals et al. · 0 citations
Review Aug 2026

Enhancing Clinical Decision-Making Using Generative AI-Powered Knowledge Retrieval Systems: A Review of Emerging Approaches and Challenges

The study concludes that evidence-based, auditable, locally adaptable, locally adaptable, and supervised by licensed clinician retrieval systems with generative AI can support safer, faster, and more relevant decision-making processes in clinical settings.

Sonam Kumari · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.