Skip to content
Review Open access

Citation reliability of frontier large language models in medical writing and its automated verification

Unknown authors
Sep 2026 · medRxiv · 0 citations
Medicine

Abstract

Large language models (LLMs) are increasingly used to draft medical manuscripts, yet their citations are unreliable and clinicians lack a validated way to verify them. We evaluated three frontier LLMs, Claude Opus 4.8, GPT-5.5, and Gemini 3.5 Flash, generating 270 cardiology narrative reviews with web search enabled, and verified all 8,050 references against PubMed. Problematic references accounted for 11.5% of GPT-5.5 output, 29.2% of Claude output, and 29.6% of Gemini output (P < 0.001), with no significant gradient across topics of differing publication volume (P = 0.052). Misattribution, a valid PubMed identifier that resolves to a different article, made up 77% of errors, whereas fabrication was rare (0.6%). Against an expert-adjudicated set of 270 references, an LLM-based Chain-of-Verification (CoVe) detected 60 of 62 problematic references (sensitivity 96.8%, specificity 98.6%), including every misattribution and fabrication. LLM-generated citations require identifier-level verification, and CoVe provides it at expert-level accuracy.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.