Skip to content
Open access

Expert-Rated Documentary Quality of AI-Assisted Hospital Discharge Reports: A Retrospective Paired Comparison with Physician-Written Reports

Jul 2026 · Healthcare · Vol 14 · 0 citations · 20 references
Medicine

TL;DR

AI-assisted discharge reports received higher expert-rated documentary quality scores in a non-blinded paired evaluation across most evaluated dimensions, supporting the need for a supervised hybrid model in which AI generates the initial draft while the clinician mandatorily validates sensitive content.

Abstract

Background/Objectives: The hospital discharge report is a critical document for care continuity that generates a substantial administrative burden for clinicians. Generative artificial intelligence (AI) offers the potential to reduce this burden while improving documentary quality. This study aims to compare, under real-world conditions with a GDPR-oriented architecture based on prior local anonymisation, the quality of AI-assisted discharge reports (IAIA) against those drafted by the responsible physician (INF). Methods: A retrospective, paired, expert-evaluation study was conducted at a Spanish university hospital. One hundred and twenty consecutive clinical cases from nine departments were included (240 reports total). Each case was independently evaluated by one of ten primary care physicians using a structured rubric covering 13 clinical dimensions (ordinal scale 1–3) and a global rating scale (1–10). The Wilcoxon signed-rank test was applied to all paired comparisons; effect size was estimated using the paired rank-biserial correlation (r). Results: IAIA achieved a significantly higher overall mean rating than INF (8.14 vs. 7.30 out of 10; p < 0.0001; r ≈ 0.76, large effect). IAIA was nominally superior in 9 of 13 clinical dimensions; after Bonferroni correction for the 13 per-dimension comparisons, six of these differences remained statistically significant, with the largest gains in family history, principal diagnosis hierarchy, and structured listing of secondary diagnoses. INF retained an advantage only in allergies and intolerances (2.69 vs. 2.46; p = 0.002), where IAIA tended to use generic formulas. Three dimensions showed no significant difference (prior treatment, physical examination, procedures). Conclusions: AI-assisted discharge reports received higher expert-rated documentary quality scores in a non-blinded paired evaluation across most evaluated dimensions. The physician-written report retained an advantage only in the safety-critical allergy domain, where allergy information must not be inferred by the model but sourced from verified structured fields or explicitly flagged as pending physician validation, supporting the need for a supervised hybrid model in which AI generates the initial draft while the clinician mandatorily validates sensitive content. Prior local anonymisation constitutes a GDPR-oriented approach to generative AI deployment in European hospital settings, substantially reducing the risk of disclosure of identifiable clinical information.

Read PDF

Similar papers

Review Aug 2026

Clinical evaluation of a vision-language model for optimizing triage and clinical workflows in critical care.

The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and patient monitoring methods.

I. Strechen, P. Krishnan, O. Kilickaya et al. · 0 citations
Review Open access Jul 2026

The High-Volume OPD Problem: Why Indian Clinical Documentation Requires a Purpose-Built Artificial Intelligence Model

Background: Ambient artificial intelligence (AI) clinical documentation platforms have demonstrated significant capacity to reduce physician documentation burden. However, existing commercial platforms are engineered for Western clinical ecosystems featuring 15-to-30-minute monolingual consultations. This narrative review evaluates the structural mismatches encountered when translating these architectures directly into the high-volume, short-duration, multilingual realities of Indian outpatient departments (OPDs). Methods: Electronic databases (PubMed, Scopus, IndMED, Google Scholar) were queried for clinical validation data, automatic speech recognition (ASR) performance metrics, and national healthcare workforce bulletins published between January 2017 and May 2026. A total of 23 core references were synthesized to map current systemic operational constraints. Results: The evaluation identified three acute structural mismatches: 1. A temporal conflict where natural language processing (NLP) architectures fail to reliably extract structured entities from compressed 1.9-to-6.9-minute Indian consultations. 2. A linguistic barrier where monolingual English ASR models suffer a 30% to 50% surge in word error rates when processing localized code-switched (Hinglish) dialogue. 3. An infrastructural disconnect due to the heterogeneous, often paper-based, electronic health record (EHR) footprint across Indian hospitals. Conclusion: Globally imported ambient documentation tools are structurally incompatible with Indian outpatient workflows. Resolving physician burnout securely requires establishing an India-native clinical AI research infrastructure optimized for short-consultation contexts and multi-language code-switched speech processing. Keywords: Ambient clinical intelligence, Clinical documentation, Outpatient department, Automatic speech recognition, Hinglish.

Anushtha Rakesh Chillure · 0 citations
Open access Mar 2026

Generating Patient Documents from Electronic Health Records Using Generative Artificial Intelligence: A Feasibility Study in a Japanese Cancer Center

Abstract Objectives Clinical documentation consumes substantial clinician time, potentially detracting from patient care. Generative artificial intelligence (AI) may support drafting discharge summaries and patient referral documents, but feasibility in non-Western-language oncology settings using real-world electronic health record (EHR) data remains insufficiently evaluated. This study assessed feasibility in a Japanese cancer hospital using an enterprise AI system. Methods Medical records from 61 consenting adult patients at Chiba Cancer Center were analyzed. Although the plan aimed at comprehensive EHR data, actual input was limited to extractable text (physician notes, nursing records); structured laboratory data and imaging, endoscopy, and pathology reports were not directly used, and existing summaries and external referrals were excluded to avoid information leakage. Data were converted to JavaScript Object Notation; GaiXer generated 31 discharge summaries and 30 referral documents. Four evaluators scored them; ≥80/100 was an exploratory threshold for draft-level practical utility. Feedback drove one refinement cycle. Results Generated documents scored approximately 60 to 70. A score ≥80 was reached by 9 of 31 discharge summaries in each evaluation; for referrals, none reached the threshold initially, whereas 5 of 30 did after refinement. Discharge summary scores did not substantially improve; referral scores did. Raw percent agreement among three nonphysician evaluators was high, although chance-corrected agreement varied. Wilcoxon signed-rank tests showed no significant change for discharge summaries ( p  = 0.866) but significant improvement for referrals ( p  = 0.006). Conclusion This feasibility study suggests AI may support drafting these documents in a secure environment using real-world Japanese EHR data, although the generated documents did not consistently reach the predefined threshold for draft-level utility. Findings should not be interpreted as demonstrating workload reduction or maximum performance under ideal data conditions. Future studies should evaluate larger datasets, multiple institutions and models, blinded evaluations, actual editing time, clinician acceptance, and workflow impact.

N. Michihata, Hiroshi Ishii, H. Tsujimura et al. · 0 citations
Open access Jul 2026

Effect of evaluation prompt strategies on LLM-as-a-judge reliability in critical care.

Bottom-up incremental scoring showed the closest alignment with human assessment in clinical AI evaluation, underscoring the need for standardised prompt architectures in clinical AI evaluation.

Jia-Yu Yan, Wing-Sum Chan, Ching-Tang Chiu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.