Skip to content
Open access

Concordance Between Clinical Practice Recommendations Generated by Generative Artificial Intelligence and the Vía RICA 2026 Enhanced Recovery Guideline: A Proof-of-Concept Study Using a Closed Evidence Corpus

Aug 2026 · Machine Learning and Knowledge Extraction · Vol 8, pp. 256 · 0 citations · 15 references

TL;DR

The results delimit the current utility of generative AI for guideline development and reveal a systematic granularity bias, producing 1–8 recommendations per bundle regardless of ground-truth size.

Abstract

Clinical practice guidelines require expert synthesis that large language models (LLMs) might partly automate, yet their ability to reproduce clinically actionable recommendations is poorly quantified. We evaluate an LLM (Claude Sonnet 4.6) against the 103 recommendations of the Spanish enhanced-recovery guideline Vía RICA 2026, grouped in 17 bundles. The model used the panel’s own closed corpus (617 documents) in a multilingual retrievalaugmented generation pipeline. Concordance was assessed twice: by optimal 1:1 bipartite matching (Hungarian) on cosine similarity, and by an LLM-as-a-judge clinical adjudicator (Claude Haiku 4.5) validated against a three-clinician panel (Fleiss’ κ = 0.538). The two schemes bracket a micro F1 of 0.61–0.69 and reveal four findings: (i) a systematic granularity bias, producing 1–8 recommendations per bundle regardless of ground-truth size; (ii) failure of cosine similarity to discriminate within narrow clinical domains; (iii) high reference-concordance precision (0.70–0.81) despite low exhaustiveness; and (iv) no transfer of the GRADE fields, evidence level agreeing no better than chance and strength systematically downgraded. An eight-fold larger retrieval budget left it intact. A corpus audit found 25 documents that formulate recommendations; excluding them lowers judged micro F1 to 0.602. The results delimit the current utility of generative AI for guideline development.

Read PDF

Similar papers

Review Open access Aug 2026

A Self-Controlled Benchmark of Retrieval-Augmented Generation for Large Language Models on Clinical Guideline Questions

Background/Objectives: Large language models (LLMs) show promise for clinical decision support, yet their accuracy in interpreting specialized medical guidelines remains uncertain. Retrieval-augmented generation (RAG) may enhance performance by grounding responses in authoritative knowledge bases. This study aimed to compare the accuracy, comprehensiveness, and safety of RAG-enhanced versus standard LLMs for answering clinical questions derived from the German S3 guideline for oral cavity carcinoma. Methods: We conducted a prospective, single-blind benchmark study evaluating six LLMs: one RAG-enhanced model (Custom GPT with guideline access), one consensus-based model (ConsensusGPT), and four standard models (DeepSeek-V3.2, Mistral Small 3.2, Qwen3-Next-80B, GPT-OSS-120B). Fifty clinical questions covering 17 guideline domains were presented to each model three times, yielding 900 evaluations. Three expert reviewers assessed responses using 5-point Likert scales for accuracy, comprehensiveness, and clarity, under a single-blind procedure, the effectiveness of which was tested by a pre-specified manipulation check. We then ran a paired within-model experiment in which each base model was queried with and without guideline access through a transparent, openly released retrieval pipeline, and scored every response with a condition-blind automated judge alongside deterministic retrieval metrics computed from the logs. Secondary outcomes included hallucination rates and guideline citation behavior. Inter-rater reliability was assessed using intraclass correlation coefficients (ICCs). Results: In a paired within-model design that held each base model fixed, adding transparent guideline retrieval improved accuracy—significantly in the three weaker open-weight models (Mistral, Qwen3, and GPT-OSS) and directionally in the already-strong DeepSeek and GPT-5 bases. Because a pre-specified blinding check found that experts could still identify retrieval-augmented answers with 98.5% accuracy, we anchored causal interpretation on measures that do not depend on the human raters, ranked by their independence: deterministic, log-derived retrieval metrics first, and then an automated, condition-blind LLM judge, whose agreement with the experts (Spearman ρ = 0.81, 95.7% within-one agreement) establishes shared calibration rather than independence from their bias. Deterministically from the retrieval logs, citation groundedness rose from 0% to 51–89% and retrieval recall@5 was 92%. On the judge, content-level hallucination fell from 42% to 4% and accuracy rose by a pooled +0.64 points (95% CI 0.47–0.80); the accuracy gain persisted after adjustment for response length (+0.48, 95% CI 0.22–0.73), which retrieval shortened rather than lengthened. The accuracy gain was large for weaker base models and small or non-significant for already-strong ones, whereas the hallucination and auditability gains were consistent across all models. The human ratings reproduced the judge’s accuracy effect (+0.61, 95% CI 0.49–0.74), and GPT-5 run through the transparent pipeline showed no significant difference from the proprietary Custom GPT (judge accuracy 4.48 vs. 4.58). Conclusions: Guideline retrieval yields a reproducible, largely base-independent improvement in the safety and auditability of LLM answers to clinical guideline questions, with accuracy gains concentrated in weaker base models. Because retrieval-augmented answers are recognizable to experts, rigorous evaluation should rely on rater-independent measures, and residual hallucination continues to require human oversight.

Andreas Vollmer, Lara Schorn, Felix Schrader et al. · 0 citations
Open access Feb 2026

A Fully Expert Human-Based Retrieval Augmented Generation (FEH-RAG) Framework: A Proof of Concept Study in Labelling Patients with Sjögren Syndrome

Background Accurate application of reference standards (RSs) is essential for correct decision-making in areas governed by such standards. Yet in real-world practice, even fully trained users often apply RSs inconsistently, due to cognitive overload, stress, or other contextual factors, generating misleading evidence. This problem is exemplified by the fact that up to 80% of board-certified hematologists mislabel patients with Sjögren syndrome (SS), a connective tissue disorder (CTD) associated with the greatest risk of lymphoproliferative disorders compared to other CTDs. Embedding RS-based consultations with full objectivity into the decision-making layers of digitalized and non-digitalized settings is critical for improving both decision-making and the quality of resulting evidence. The aim of this proof-of-concept (POC) study is to apply a newly developed framework, the Fully Expert Human-based Retrieval Augmented Generation (FEH-RAG) framework, to provide such a foundation for augmenting SS case labeling in routine daily clinical practice. Methods In this POC study, using the nine steps of the FEH-RAG framework, seven expert end-users systematically selected the most widely used SS classification criteria (SSCC) and extracted their elements, including items, definitions, item weights, and inter-item relationships. Extracted items were profiled based on their usage in routine clinical practice. A pathway layout and decision tables were developed accordingly. Following pathway generation, the residual misalignment of the outputs with the SSCC was assessed in a cohort of patients at risk of SS. The experts predefined that the residual misalignment rate of the FEH-RAG outputs with the SSCC must be ≤2% (95% confidence, using the rule of three). Results The FEH-RAG framework objectively generated RS-based, transparent, and traceable outputs, including decision tables, an SS classification pathway, and a list of misinterpretations of the SSCC. These misinterpretations involved definitions of dry eye and dry mouth, application of secondary SS criteria, handling SS criteria-specific exclusion rules, and interpretation of serological and objective test results. This POC study achieved its expert-defined maximum misalignment threshold of ≤2% with 95% confidence (0 misalignment in 150 consecutive patients at risk of SS). Conclusion This POC study established the needed foundation for improving SS case labeling in daily clinical practice across both digital and non-digital settings. As shown here, publishing FEH-RAG outputs while highlighting potential RS misinterpretations offers a transparent and traceable basis for augmenting decision-making in domains governed by RSs.

A. Hajiabbasi, Ahmadreza Jamshidi, S. Shoaee et al. · 0 citations
Review Open access Aug 2026

Explainability of decoder-only clinical large language models: A scoping review

Findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited, and three regulatory priorities are highlighted: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.

Nishant Mishra, Ameen Abu-Hanna, Iacer Calixto · 1 citation
Open access Jul 2026

Effect of evaluation prompt strategies on LLM-as-a-judge reliability in critical care.

Bottom-up incremental scoring showed the closest alignment with human assessment in clinical AI evaluation, underscoring the need for standardised prompt architectures in clinical AI evaluation.

Jia-Yu Yan, Wing-Sum Chan, Ching-Tang Chiu et al. · 0 citations
Open access Aug 2026

Role prompting modulates linguistic style but not clinical decision structure in GPT-5 tumour board simulation

GPT-5 reliably adapts linguistic style to clinical personas but produces limited specialty-specific output diversity, supporting its role as a decision-support adjunct rather than an autonomous specialist simulator.

Derna Stifini, Andrea Della Penna, André L. Mihaljevic et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.