Guideline concordance of large language models for ERAS colorectal surgery recommendations: a blinded, clinician-rated comparison of Google Gemini and ChatGPT
Jun 2026· Cukurova Anestezi ve Cerrahi Bilimler Dergisi· 0 citations· 9 references
TL;DR
Both LLMs showed high overall concordance with ERAS colorectal recommendations, however, safety‑relevant deviations were not rare, and rater agreement varied by model, highlighting the need for clinician oversight and standardized evaluation frameworks before bedside use.
Abstract
Background:Large language models (LLMs) are increasingly used by clinicians and trainees for perioperative decision support, yet their alignment with Enhanced Recovery After Surgery (ERAS) recommendations remains uncertain.Methods:We converted the 2025 ERAS Society recommendations for elective colorectal surgery into a 52‑question bank (preoperative (n=28), intraoperative (n=10), and postoperative (n=14). Each question was asked once, in a new chat, to Google Gemini and OpenAI ChatGPT (web interfaces; no follow‑up prompts or regeneration; queries performed on 17 Feb 2026). Responses were blinded as A/B and independently scored by two clinicians (an anesthesiologist and a general surgeon) for (i) guideline concordance on a 5‑point Likert scale and (ii) safety risk (0=none, 1=potential harm, 2=critical harm). Primary analysis used paired Wilcoxon signed‑rank tests (Likert) and exact McNemar tests (any safety flag ≥1). Inter‑rater agreement was estimated with quadratic weighted kappa (QWK).Results:A total of 208 ratings were generated (52 questions × 2 models × 2 raters). Mean Likert concordance was 4.49±0.62 for Gemini and 4.44±0.65 for ChatGPT (paired Wilcoxon p=0.604). Any safety flag occurred in 20.2% (Gemini) and 15.4% (ChatGPT) of ratings (McNemar p=0.383); no responses were rated as critical harm. Inter‑rater agreement was lower for Gemini (QWK=0.225) than for ChatGPT (QWK=0.597).Conclusions:Both LLMs showed high overall concordance with ERAS colorectal recommendations, with no significant overall difference in scores. However, safety‑relevant deviations were not rare, and rater agreement varied by model, highlighting the need for clinician oversight and standardized evaluation frameworks before bedside use.Keywords:ERAS; colorectal surgery; large language model; ChatGPT; Gemini;
Background/Objectives: As patients increasingly rely on large language models (LLMs) for Chronic Rhinosinusitis (CRS) diagnosis, surgical candidacy, and perioperative care, evaluating the accuracy of LLM-generated information against established clinical practice guidelines for surgical management of CRS is essential. Methods: ChatGPT, Google AI, Google Gemini, and Grok were queried using a 21-question guideline-mapped prompt set (long) and a single patient-focused prompt (short). Two physician reviewers independently scored responses using a 3-point rubric across 21 fields. Primary outcomes were guideline-concordant scores; secondary outcomes included readability measured with the Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Inter-rater reliability (IRR) was assessed using the intraclass correlation coefficient (ICC). Analyses were performed in SPSSv31. Results: Guideline concordance ranged from 55.36% to 77.98% (p > 0.05), highest for Grok (77.98%, 95% CI 63.67–92.28), followed by Google Gemini (66.67%, 95% CI 35.42–97.91), ChatGPT (55.95%, 95% CI 29.25–82.65), and Google AI (55.36%, 95% CI 29.97–80.75), with Grok significantly outperforming both ChatGPT and Google AI. Prompt structure significantly affected scores. Long-form prompting resulted in higher guideline concordance scores than short-form prompting (+26.19, p < 0.001). The CPG demonstrated a more readable structure, with a higher FRE (44.1), exceeding scores generated by Grok (31.7), ChatGPT (39.9), Gemini (39.2), and Google AI (32.5). In contrast, the CPG was a higher reading grade level (FKGL score of 11.7) than Grok (11.4), Gemini (10.5), and ChatGPT (10.3), but was lower in reading grade compared to Google AI, which produced the highest FKGL score (12.5). IRR was high (ICC = 0.961). Conclusions: LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms. While LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability. However, longer prompt structure meaningfully influenced output quality, highlighting how a user’s ability to frame precise prompts is critical to obtaining accurate information.
Hetal Lad, Emily S Kwon, Ayushi Chadha et al.· Journal of Otorhinolaryngolo...· 0 citations
INTRODUCTION
The management of rectal adenocarcinoma requires navigation of complex, branching guideline pathways encompassing neoadjuvant sequencing, surgical approach, organ preservation, and surveillance, yet real-world guideline adherence remains as low as 60-70%. The ability of current-generation large language models (LLMs) to accurately navigate these decision points has not been fully characterized.
METHODS
In this cross-sectional, vignette-based study, 135 clinical questions were constructed from 45 pages of NCCN Rectal Cancer Guidelines (Version 4.2024). ChatGPT-4o was queried using standardized prompts with up to 3 clarifying questions permitted per query. Responses were independently evaluated by two physician raters on a 5-point Likert scale, with potential discrepancies adjudicated by a board-certified surgical oncologist. Primary outcomes were the proportion of responses rated Correct (score ≥ 3) and Accurate (score ≥ 4). Inter-rater reliability was assessed using Cohen's kappa, and subgroup analysis was performed across clinical domains using the Kruskal-Wallis test.
RESULTS
Of 135 questions, 127 (94.1%; 95% CI, 88.7-97.0%) were Correct and 121 (89.6%; 95% CI, 83.3-93.7%) were Accurate. One hundred two responses (75.6%) were completely correct without additional prompting. Performance was consistent across clinical domains (Kruskal-Wallis H = 0.530, p = 0.767). Inter-rater agreement was perfect (κ = 1.0). Eight responses (5.9%) contained partially or wholly incorrect information, with errors concentrated in multi-step conditional treatment decision points.
CONCLUSION
ChatGPT-4o demonstrates high concordance with NCCN rectal cancer guidelines across all evaluated clinical domains with notable improvement over prior ChatGPT iterations evaluated by our group. The concentration of errors in complex conditional treatment algorithms suggests that LLMs excel at discrete factual recall but may struggle with multi-step reasoning under clinical uncertainty. Prospective validation using real-world clinical data and comparison with multidisciplinary tumor board recommendations remain necessary prior to clinical integration.
Ryan J Meyer, Tamir E. Bresler, Kevin Palmer et al.· Journal of Surgical Oncology· 0 citations
6
Background:
Clinical trial participation in oncology is frequently hindered by complex, jargon-laden information that impairs comprehension by patients and caregivers. LLMs offer a scalable solution to simplify technical text, but their utility in oncology has not been evaluated.
Methods:
We conducted a randomised, controlled, three-period crossover study at two tertiary breast cancer outpatient clinics in Sydney, Australia. Eligible patients and caregivers evaluated trial descriptions across 3 formats: standard ClinicalTrials.gov text (Control); LLM-optimised text generated using zero-shot prompting with GPT-4 (LLM); and further refined by an oncologist (LLM+E). Latin square randomisation controlled for order effects. The primary endpoint was the Global Preference Score (GPS), operationalised as the minimum Likert rating (1-5 scale) across 5 sections (Title, Summary, Intervention, Description, Eligibility) to reflect that incomprehensibility of any component degrades the overall document utility. Sample size (n≥18) was determined via Monte Carlo simulation to detect a minimum 1 point Likert shift with 90% power (α=0.05); the recruitment target was 36 (accounting for 50% attrition). Primary analysis employed Friedman rank sum test blocked by participant, with a cumulative link mixed model (CLMM) including random intercepts for participants to adjust for age, education, and first language. All LLM generated text (LLM±E) were vetted for accuracy of content by ≥1 medical oncologist.
Results:
Between September and December 2025, 30 of 31 recruited participants provided crossover responses for primary analysis (401 valid responses across 5 sections, 11% invalid). The mean age was 53 years (95% CI, 44-63), with the majority of respondents being female (n = 27, 90%), native English speakers (n = 24, 80%) and holders of tertiary qualifications (n = 17, 57%); 20 respondents were patients (67%). The median GPS was 3.0 (IQR 2.5) for Control, 4.0 (IQR 2.0) for LLM and 3.0 (IQR 1.5) for LLM+E. LLM achieved a significantly higher GPS vs. Control (p = 0.02, Friedman test). Multivariable CLMM confirmed increased odds of higher preference ratings for LLM text vs. Control (OR 1.83, 95% CI 1.01–3.35, adjusted p = 0.048). Paradoxically, LLM+E showed no improvement over Control (padj = 0.826, r = 0.063). Native English speakers rated contents more critically than non-native speakers (OR 0.092, p = 0.009), independent of arm allocation.
Conclusions:
LLM optimisation improves patient preference for trial information compared with standard registry descriptors, supporting further evaluation of its use in rendering patient-facing materials. Expert oncologist revision may inadvertently re-introduce complexity, negating the linguistic accessibility gains provided by LLMs, suggesting that human-in-the-loop workflows need caution to preserve linguistic style for accessibility.
M. Tran, Kate Saw, Jeremy Mo et al.· Journal of Clinical Oncology· 0 citations
INTRODUCTION
This study evaluated concordance between treatment-pathway outputs generated from structured clinical text by the large language model (LLM) GPT-5.4 Thinking and multidisciplinary team (MDT) decisions for mid- and low rectal cancer.
MATERIALS AND METHODS
This single-center retrospective study included 260 patients who underwent standardized assessment, MDT discussion, and curative-intent surgery between January 2022 and December 2023. After database lock, structured de-identified clinical text was entered into GPT-5.4 Thinking using prespecified preoperative and postoperative templates, with MDT decisions treated as real-world reference decisions. The primary endpoint was preoperative concordance; secondary endpoints included postoperative and overall concordance. Concordance metrics, directional discordance, baseline comparators, repeatability in a 50-case subset, and exploratory logistic regression models were assessed.
RESULTS
Preoperative, postoperative, and overall concordance rates were 57.7%, 66.9%, and 46.9%, respectively; Cohen's κ values were 0.269 and 0.449 for the preoperative and postoperative stages. Preoperatively, directional discordance relative to MDT decisions included 48 potential under-intensification, 28 potential over-intensification, and 34 directionally indeterminate or heterogeneous cases. Compared with the majority-class baseline, the LLM had the same preoperative crude concordance but higher κ and balanced category-specific concordance; postoperatively, it exceeded the majority-class baseline across these metrics. Non-identical mapped categories across three runs occurred in 12/50 preoperative and 8/50 postoperative assessments despite identical inputs.
CONCLUSION
GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited. Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs.
Yuanze Wei, Yulong Tian, Xiaodong Liu et al.· European Journal of Surgical...· 0 citations