Skip to content
Review Open access

Large Language Model versus Clinician Written Summaries of Research Papers.

Unknown authors
Sep 2026 · Journal of the American Board of Family Medicine · Vol 39 1 · 0 citations
Medicine

Abstract

INTRODUCTION Clinicians require concise, accurate summaries of new research to inform practice. Patient-Oriented Evidence that Matters (POEMs), published in American Family Physician, are a benchmark for summarizing primary literature in family medicine, while large language models (LLMs) offer scalable summarization but require rigorous evaluation. The objective of this study was to evaluate the accuracy and quality of summaries generated by large language models compared with expert-authored POEMs.

Methods

In this study, we compared LLM-generated summaries (Microsoft Copilot, GPT-4o class) with 24 recent matched POEMs using a standardized prompt. Two trained raters independently scored each summary with a 13-item tool (score range 0-13), cataloged errors, recorded word counts, and indicated preferences on a 5-point scale.

Results

LLM summaries outperformed POEMs in total score (mean 12.1 vs 10.6; mean difference 1.5, 95% CI 1.1-2.0; P < 0.001), with similar lengths (328 vs 353 words; P = 0.23). Errors occurred in fewer LLM-DOCSs (2/24) than POEMs (9/24), with a mean error score difference of 20% (95% CI 7% -33%; P < 0.001). POEMs most often missed in the categories Contextual Background and Limitations; both approaches frequently missed in Clinical Applicability. Reviewer preference favored LLM-DOCS (mean 2.44 on a 1-5 scale; 95% CI 2.1-2.8).

Conclusions

An enterprise LLM, prompted in POEM style, produced accurate, low-error clinical summaries that matched or exceeded expert-edited POEMs and were generally preferred by reviewers, though further research is needed to assess broader applicability and impact. Findings support pragmatic LLM-assisted summarization and highlight the need for standardized evaluation tools and explicit prompts for clinical applicability.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.