Jul 2026· World Journal of Surgery· 0 citations· 5 references
Medicine
TL;DR
In this scenario-based early evaluation, GPT-5 was the most safety-aligned model, GPT-4 was the most operationally actionable, and Gemini was strongest for synthesis, but all require clinician oversight, prospective validation, and governance before clinical deployment.
Abstract
Background
Early decision-making in acute pancreatitis (AP) involves diagnostic confirmation, early severity triage, escalation thresholds, and initiation of guideline-concordant management under time pressure and incomplete information. Large language models (LLMs) may support structured bedside reasoning, but their clinical usefulness cannot be inferred from guideline knowledge alone.
Methods
A cross-sectional, scenario-based comparative evaluation was conducted in January 2026 using 20 AP scenarios: 15 refined hypothetical vignettes and 5 de-identified, privacy-modified real-life case patterns. GPT-4, GPT-5, and Gemini received identical single-turn prompts. Model access was through OpenAI API gpt-4-0613, OpenAI API gpt-5, and Google Vertex AI Gemini 1.0 Pro; temperature was set to 0.0, and each prompt was repeated three times per model. Outputs were scored by two independent clinician-raters using a prespecified 1-5 ordinal rubric across guideline concordance, safety, actionability, and data-synthesis quality. Two senior board-certified surgeons independently generated expert reference pathways for comparison.
Results
GPT-5 achieved the highest guideline concordance (4.28 ± 0.38) and safety (4.20 ± 0.45) profiles. GPT-4 provided the clearest stepwise actionability (4.15 ± 0.48), whereas Gemini showed the strongest data-synthesis quality (4.22 ± 0.52). With deterministic settings, internal consistency across three repeated runs was 100%. All models demonstrated clinically relevant failure modes, particularly unwarranted certainty under missing data; this occurred in 12/20 GPT-4, 7/20 GPT-5, and 15/20 Gemini outputs.
Conclusion
No model should be used as a stand-alone bedside decision-maker for AP. In this scenario-based early evaluation, GPT-5 was the most safety-aligned model, GPT-4 was the most operationally actionable, and Gemini was strongest for synthesis, but all require clinician oversight, prospective validation, and governance before clinical deployment.
BACKGROUND
Clinical decision-making requires integrating history, physical examination, laboratory, and imaging data. In the emergency department (ED), workload, time pressure, and cognitive burden may impair this process and affect decision quality. This study compares the diagnostic outputs of ChatGPT, Claude, and Gemini with those of emergency physicians in real-world ED cases.
METHODS
This prospective, single-centre observational diagnostic agreement study compared the stage-wise outputs of four Large Language Models (LLMs) (ChatGPT-4o, ChatGPT-5, Claude Opus 4.1, and Gemini 2.5 Pro) with those of emergency physicians in critically ill ED patients. Between 10 August and 10 September 2025, de-identified clinical data were entered into the models via their official web interfaces using standardised prompts. In the first stage, physicians and LLMs each generated five preliminary diagnoses based on vital signs and medical history. In the second stage, following physical examination and laboratory and imaging results, both refined their lists into three differential diagnoses. In the third stage, the physicians' final diagnosis was accepted as the reference, and each LLM was prompted to provide a final diagnosis. LLM preliminary and differential diagnoses were compared with those of the physicians at the corresponding stage, and LLM final diagnoses with the reference; the inclusion of the final diagnosis within earlier lists was also evaluated. Agreement was quantified using Cohen's κ; analyses were performed in R.
RESULTS
Of 389 screened patients, 180 were included (56.1% male; mean age 67 ± 15.9 years). Physicians contained the reference diagnosis within their top-5 preliminary and top-3 differential lists in 83.9% and 98.3% of cases, respectively, significantly exceeding every LLM (all p < 0.001). Final-diagnosis match rates were 67.2% [60.3-73.5] for ChatGPT-4o, 65.6% [58.7-71.9] for ChatGPT-5, 63.3% [56.3-69.9] for Claude Opus 4.1, and 59.4% [52.3-66.1] for Gemini 2.5 Pro (p = 0.16). Cohen's κ ranged from 0.575 (Gemini 2.5 Pro) to 0.656 (ChatGPT-4o), indicating moderate-to-substantial agreement, with no pairwise difference reaching significance.
CONCLUSIONS
The LLMs achieved moderate agreement with ED reference diagnoses in critically ill patients but were consistently outperformed by physicians at the early diagnostic phases. Despite final-diagnosis match rates of 59%-67%, their current diagnostic role in the ED remains limited.
İbrahim Günaydın, M. Yılmaz, Sinan Akpunar et al.· BMC Emergency Medicine· 0 citations
Large language models outperformed emergency medicine physicians in overall diagnostic accuracy and demonstrated superior consistency across different types of diagnostic tasks, indicating that AI possesses robust pattern-recognition and reasoning capabilities, suggesting it could serve as a highly reliable clinical decision support tool.
Mehdi Arzani Shamsabadi, Roya Vatankhah, Hasan Jalilvand et al.· International Journal of Eme...· 0 citations
Background Patients increasingly rely on large language models (LLMs) for health information, yet their suitability for decision-critical conditions such as acute pancreatitis remains unclear. Given that acute pancreatitis requires timely symptom recognition, severity assessment, treatment decision-making, recurrence prevention, and follow-up management, LLM-generated information should demonstrate reliability, transparency, and readability. Objectives To evaluate the informational quality, visible transparency-related features, and readability of English-language responses generated by five publicly accessible LLMs to standardized patient-facing questions on acute pancreatitis. Methods This cross-sectional benchmark study developed 24 English-language, single-intent questions on acute pancreatitis across six clinical domains using public search intents and guideline-derived decision-critical content. The analysis focused exclusively on English-language patient-facing responses. Each question was submitted once to GPT-5.4 Thinking, DeepSeek-V3.2, Gemini 3.1 Pro, Grok 4.3, and Qwen3.6-Max-Preview, generating 120 responses. Anonymized responses were independently evaluated by two blinded gastroenterologists using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale (GQS), and Journal of the American Medical Association (JAMA) benchmark criteria. Readability was assessed using six established formulas. A structured response-level safety analysis was added to evaluate factual inaccuracies, clinical hallucinations, and clinical safety-risk severity. Between-model differences were analyzed with Friedman tests followed by post hoc paired Wilcoxon signed-rank tests with Holm correction. Results Significant between-model differences were observed across all quality, transparency-related, and readability outcomes. Grok 4.3 achieved the highest mean DISCERN, GQS, EQIP, and JAMA scores, reflecting the strongest informational quality profile according to the predefined quality instruments, although visible transparency cues remained limited across all models. In the added safety analysis, factual inaccuracies and clinical hallucinations were each identified in 16 of 120 responses (13.3%), whereas clinical safety-risk signals were identified in 7 responses (5.8%), all of which were adjudicated as score 1 (low risk) under the predefined clinician-rated rubric; no moderate- or high-risk clinical safety event was adjudicated. DeepSeek-V3.2 demonstrated the most favorable readability profile, with the highest Flesch Reading Ease score (44.42 ± 12.27), which nevertheless remained substantially below the recommended threshold of ≥ 80. None of the 120 responses satisfied all six predefined readability targets. All model-level readability distributions differed significantly from recommended thresholds in the direction of poorer readability. Conclusion Publicly accessible LLMs generated English-language responses on acute pancreatitis with variable informational quality, limited visible transparency cues, and consistently inadequate readability under standardized default public-interface conditions. Because the primary analysis used a single-generation cross-sectional design, model rankings should be interpreted as performance snapshots rather than definitive or temporally stable hierarchies. Higher quality scores did not necessarily establish factual accuracy, clinical safety, patient-education-level readability, or stronger response-level transparency. Current public-interface LLM outputs may serve as clinician-reviewed drafts for patient education but should not function as standalone patient resources. Future AI-based health information systems should strengthen clinical completeness, plain-language communication, visible evidence support, actionability, and clinician-supervised safeguards.
Biao Jiang, Hong-Xin Sun, Linlin Chen· Frontiers in Public Health· 0 citations
Standardizing emergency admission decisions for acute diverticulitis (AD) is essential to optimize bed utilization and patient safety. This study evaluates a AI based machine learning (ML)–based model designed to categorize AD severity and recommend admission versus outpatient management by integrating structured and free-text data.
Data of all patients admitted with AD in specified time period was analysed using a hybrid natural language processing (NLP) and ML framework. The model utilized 14 key parameters, including demographics (age), physiology (HR, SBP, Temp, GCS), and biochemistry (CRP, WCC, Creatinine, INR). Regex-based NLP extracted clinical features from triage notes, and CT reports, specifically identifying LIF pain, and immunosuppression. Severity grading and admission logic were aligned with NICE, ACPGBI, and WSES guidelines, incorporating specific thresholds such as CRP>150 (warning level) and >200 mg/L (emergency level) to identify complicated disease. A logistic regression classifier with balanced weights was trained and evaluated using an 80/20 train-test split.
The model successfully stratified patients into three severity grades. Grade 3 (Severe) was triggered by organ dysfunction or complications (perforation/abscess), while Grade 2 (Moderate) identified high-risk factors like age>65, ASA⩾3, or immunosuppression. The ML classifier demonstrated high predictive performance, achieving test-set ROC AUC of 1.0, indicating perfect discrimination between the rule-based admission recommendations and clinical features within this pilot cohort.
By integrating structured biochemical data with NLP-extracted risk factors, this model provides a reliable clinical “second opinion.” It demonstrates significant potential for enhancing triage accuracy and ensuring guideline-compliant management of acute diverticulitis.
Fatima Zulfiqar, Yasser Mohamed, Raghav Anirudh et al.· British Journal of Surgery· 0 citations
On complex ID scenarios, large language models responses were variable and caution is required when deploying these models in ID domains without specialist oversight, suggesting caution is required when deploying these models in ID domains without specialist oversight.
A. Pradhan, B. Waxse, W. Matias et al.· medRxiv· 0 citations
Domain-specific medical LLMs enhanced by CoT approach the antibiotic decision-making level of real physicians, with advantages in individualization and dosing precision, but notable deficiencies persist in antimicrobial stewardship ecological awareness and automated evaluation reliability, underscoring the continued indispensability of senior clinical expertise.
Y. Liu, C. Zhang, F. Wang et al.· medRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.