These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk, and show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.
Abstract
Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models'training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.
Large language models (LLMs) are increasingly used to draft medical manuscripts, yet their citations are unreliable and clinicians lack a validated way to verify them. We evaluated three frontier LLMs, Claude Opus 4.8, GPT-5.5, and Gemini 3.5 Flash, generating 270 cardiology narrative reviews with web search enabled, and verified all 8,050 references against PubMed. Problematic references accounted for 11.5% of GPT-5.5 output, 29.2% of Claude output, and 29.6% of Gemini output (P < 0.001), with no significant gradient across topics of differing publication volume (P = 0.052). Misattribution, a valid PubMed identifier that resolves to a different article, made up 77% of errors, whereas fabrication was rare (0.6%). Against an expert-adjudicated set of 270 references, an LLM-based Chain-of-Verification (CoVe) detected 60 of 62 problematic references (sensitivity 96.8%, specificity 98.6%), including every misattribution and fabrication. LLM-generated citations require identifier-level verification, and CoVe provides it at expert-level accuracy.
Findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited, and three regulatory priorities are highlighted: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.
INTRODUCTION
Clinicians require concise, accurate summaries of new research to inform practice. Patient-Oriented Evidence that Matters (POEMs), published in American Family Physician, are a benchmark for summarizing primary literature in family medicine, while large language models (LLMs) offer scalable summarization but require rigorous evaluation. The objective of this study was to evaluate the accuracy and quality of summaries generated by large language models compared with expert-authored POEMs.
METHODS
In this study, we compared LLM-generated summaries (Microsoft Copilot, GPT-4o class) with 24 recent matched POEMs using a standardized prompt. Two trained raters independently scored each summary with a 13-item tool (score range 0-13), cataloged errors, recorded word counts, and indicated preferences on a 5-point scale.
RESULTS
LLM summaries outperformed POEMs in total score (mean 12.1 vs 10.6; mean difference 1.5, 95% CI 1.1-2.0; P < 0.001), with similar lengths (328 vs 353 words; P = 0.23). Errors occurred in fewer LLM-DOCSs (2/24) than POEMs (9/24), with a mean error score difference of 20% (95% CI 7% -33%; P < 0.001). POEMs most often missed in the categories Contextual Background and Limitations; both approaches frequently missed in Clinical Applicability. Reviewer preference favored LLM-DOCS (mean 2.44 on a 1-5 scale; 95% CI 2.1-2.8).
CONCLUSIONS
An enterprise LLM, prompted in POEM style, produced accurate, low-error clinical summaries that matched or exceeded expert-edited POEMs and were generally preferred by reviewers, though further research is needed to assess broader applicability and impact. Findings support pragmatic LLM-assisted summarization and highlight the need for standardized evaluation tools and explicit prompts for clinical applicability.
Unknown authors· Journal of the American Boar...· 0 citations
Current evidence supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.
Yun-Jia Wu, Qi Yan, Dingcheng Tian· AI Medicine· 0 citations
Five widely used prompting strategies across five influential LLMs in the latest medical bias benchmark reveal substantial heterogeneity in both effectiveness and overhead across models, with no strategy proving universally effective and some even exacerbating bias.
Ying Xiao, Zhenpeng Chen, Jie M. Zhang· Philosophical transactions....· 1 citation
Clinical AI evaluation should encompass diagnosis and management after adaptive information gathering. We compared Doctorina, eight physicians and four standalone frontier language models in 150 synthetic Polish-language primary-care consultations. Doctorina achieved 82.0% Top-1 concordance versus 57.0% for physicians (difference, 25.0 percentage points; 95% confidence interval, 17.7-32.7) and 97.3% versus 85.0% primary-or-reference-differential concordance. Across 149 case pairs, normalized workup and treatment scores were 89.4 versus 66.9 and 83.7 versus 61.2. Doctorina had the highest diagnostic point estimates among all six groups; Kimi K3 ranked next, while Claude Opus 5 led the closely spaced management estimates of Opus, Doctorina and Kimi. A second Doctorina execution reproduced the advantages over physicians across all outcomes. Doctorina's advantage over physicians therefore extended from primary-diagnosis selection to higher-rated diagnostic workup and initial treatment after adaptive consultation.
Andy Nkansah, Hanna Plotnitskaya, Stanislau Salavei et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.