Skip to content
Open access

Comparison of physician-authored and artificial intelligence-generated after-visit summaries: A blinded comparative study.

Sep 2026 · Journal of Hospital Medicine · 0 citations · 21 references
Medicine

TL;DR

In this blinded evaluation, LLM-generated AVSs were clearer, more comprehensive, and more empathetic than physician-authored AVSs and were associated with lower physician-rated potential for harm.

Abstract

Background

After-visit summaries (AVSs) are essential for a safe hospital discharge, yet are often written above literacy levels, omit key information, and are produced under substantial clinical time pressure. Large language models (LLMs) offer a potential solution, while performance and safety in real clinical workflows remain uncertain.

Objective

To compare the quality, understandability, actionability, and safety of LLM-generated versus physician-authored after-visit summaries for hospitalized patients.

Methods

We conducted a retrospective, blinded comparison of physician-authored AVSs and LLM-generated AVSs for 50 adults discharged from the University of California San Diego Health hospital medicine service in 2023. The physician-authored hospital course served as source text for generating AVSs using GPT-4 and Gemma 3n 2. Both were prompted to produce sixth-grade level, patient-centered AVSs. Five attending physicians independently evaluated each AVS using the Patient Education Materials Assessment Tool (PEMAT) for understandability and actionability and the AVSrubric, an instrument assessing accuracy, comprehensiveness, clarity, consistency with the medical record, tone and empathy, and potential for harm.

Results

LLM-generated AVSs had higher PEMAT scores than physician-authored AVSs (understandability: 85.5% (GPT-4), 87.5% (Gemma), and 66.1% (physician-authored); actionability: 70.9% (GPT-4), 74.1% (Gemma) vs. 56.7% (physician-authored); all p < .001). On the AVSrubric, LLM-generated AVSs received higher ratings than physician-authored AVSs across all domains with GPT-4 demonstrating significantly higher in all five, while Gemma showed significant improvements in clarity, readability, tone, and empathy.

Conclusions

In this blinded evaluation, LLM-generated AVSs were clearer, more comprehensive, and more empathetic than physician-authored AVSs and were associated with lower physician-rated potential for harm. A physician-in-the-loop LLM workflow may improve discharge communication while reducing clinician burden and warrants prospective evaluation.

Read PDF

Similar papers

#large language models Review Sep 2026

A Real-World Evaluation of Large Language Model-Generated Hospital Courses in Pediatrics.

BACKGROUND Large language model (LLM)-generated hospital courses are increasingly integrated into electronic health records (EHRs), yet their accuracy and safety in pediatric populations remain poorly characterized. OBJECTIVE To evaluate the accuracy, text quality, and perceived potential harm of EHR-integrated and L...

Jasmine E. Kim, J. Hron, Daniel J Kats et al. · 0 citations
Review Open access Aug 2026

Readability of AI-Generated Patient Visit Summaries in Orthopedic Surgery: Retrospective Analysis

Abstract Background Patient visit summaries (PVS) are patient-facing documents intended to reinforce communication and promote patient education after clinical encounters. Despite national recommendations that patient education materials be written at or below a sixth-grade reading level, most orthopedic materials subs...

E. L. Major, Vivek P. Shah, Amber N. Carroll et al. · 0 citations
Review Open access Sep 2026

Large Language Model versus Clinician Written Summaries of Research Papers.

An enterprise LLM, prompted in POEM style, produced accurate, low-error clinical summaries that matched or exceeded expert-edited POEMs and were generally preferred by reviewers, though further research is needed to assess broader applicability and impact.

Richard Guthmann, Robert Martin, Erin Lee et al. · 0 citations
Review Sep 2026

A comparative evaluation of LLM-generated and clinician-written psychotherapy session summaries.

OBJECTIVE Psychotherapy session summaries are crucial to the continuity of care, supervision, and treatment planning; however, they are time-consuming and contribute to the documentation burden. Large language models (LLMs) may help streamline this workflow, but their adequacy must be established. We evaluated whether...

Liran Keren, Ayal Klein, Yael Bar-Shachar et al. · 0 citations
Review Open access Aug 2026

Physician-revised AI-generated drafts are associated with higher ratings of written explanations in end-of-life care in the intensive care unit: a scenario-based single-center cross-sectional study

In end-of-life care (EOL) in the intensive care unit (ICU), intensivists are expected to provide medically appropriate and empathetic communication to support shared decision-making with patients and their families. Large language models (LLMs) have shown potential to generate medical responses that are perceived as in...

Atsushi Kokita, Junpei Haruna, Yuya Goto et al. · 0 citations
Aug 2026

Quality of AI-Generated Patient Education for Pre- and Post-Operative Tracheostomy Care.

AI chatbots can generate accurate and comprehensive responses to common tracheostomy care questions, demonstrating potential to support patient education, but they continue to lack guaranteed, verifiable sourcing.

Keer Zhang, Lauran K. Evans, Desiree Delavary et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.