Skip to content

Clinical safety of large language model responses to matched patient-language and clinician-language Turkish obstetric and gynecologic triage prompts: a model-blinded paired-scenario study.

Aug 2026 · European Journal of Obstetrics, Gynecology, and Reproductive Biology · Vol 326, pp. 115351 · 0 citations · 19 references
Medicine

Abstract

Objective

To determine whether presenting matched obstetric and gynecologic triage scenarios as patient-language prompts rather than clinician-language prompts affects expert-rated clinical confidence and the clinical safety of LLM-generated advice.

Methods

Thirty obstetric and gynecologic scenarios were presented in matched clinician-language and patient-language Turkish formats to four LLMs. Five specialists independently evaluated 240 responses, generating 1,200 ratings. The primary outcome was the Global Clinical Confidence Score (GCCS; 0-2); five secondary outcomes were rated on 1-5 scales. Associations were examined using ordinal logistic generalized estimating equations adjusted for model and evaluator.

Results

Clinically reliable responses (GCCS = 2) accounted for 89.3% of clinician-language and 91.3% of patient-language ratings. Patient-language phrasing was not significantly associated with overall GCCS (cumulative odds ratio 0.78, 95% confidence interval 0.57-1.06; p = 0.115), and the language-by-model interaction was not significant (p = 0.422). Patient-language prompts were associated with fewer GCCS = 0 ratings in a binary generalized estimating equations analysis (odds ratio 0.66, 95% confidence interval 0.46-0.96; p = 0.031), although the exact paired McNemar test was not significant (p = 0.096). After false-discovery-rate correction, patient-language prompts had higher evaluator-level triage appropriateness and clinical applicability scores (both adjusted p = 0.028).

Conclusion

No significant difference in overall expert-rated clinical confidence was detected between patient-language and clinician-language prompts.

View source

Similar papers

Review Open access Aug 2026

Benchmarking Large Language Model Responses Against Surgical Clinical Practice Guidelines for Chronic Rhinosinusitis: The Importance of User Prompts

Background/Objectives: As patients increasingly rely on large language models (LLMs) for Chronic Rhinosinusitis (CRS) diagnosis, surgical candidacy, and perioperative care, evaluating the accuracy of LLM-generated information against established clinical practice guidelines for surgical management of CRS is essential. Methods: ChatGPT, Google AI, Google Gemini, and Grok were queried using a 21-question guideline-mapped prompt set (long) and a single patient-focused prompt (short). Two physician reviewers independently scored responses using a 3-point rubric across 21 fields. Primary outcomes were guideline-concordant scores; secondary outcomes included readability measured with the Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Inter-rater reliability (IRR) was assessed using the intraclass correlation coefficient (ICC). Analyses were performed in SPSSv31. Results: Guideline concordance ranged from 55.36% to 77.98% (p > 0.05), highest for Grok (77.98%, 95% CI 63.67–92.28), followed by Google Gemini (66.67%, 95% CI 35.42–97.91), ChatGPT (55.95%, 95% CI 29.25–82.65), and Google AI (55.36%, 95% CI 29.97–80.75), with Grok significantly outperforming both ChatGPT and Google AI. Prompt structure significantly affected scores. Long-form prompting resulted in higher guideline concordance scores than short-form prompting (+26.19, p < 0.001). The CPG demonstrated a more readable structure, with a higher FRE (44.1), exceeding scores generated by Grok (31.7), ChatGPT (39.9), Gemini (39.2), and Google AI (32.5). In contrast, the CPG was a higher reading grade level (FKGL score of 11.7) than Grok (11.4), Gemini (10.5), and ChatGPT (10.3), but was lower in reading grade compared to Google AI, which produced the highest FKGL score (12.5). IRR was high (ICC = 0.961). Conclusions: LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms. While LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability. However, longer prompt structure meaningfully influenced output quality, highlighting how a user’s ability to frame precise prompts is critical to obtaining accurate information.

Hetal Lad, Emily S Kwon, Ayushi Chadha et al. · 0 citations
Jul 2026

The illusion of competence: Evaluating the clinical reasoning of large language models in pediatric gastroenterology.

OBJECTIVE To evaluate the diagnostic accuracy and clinical reasoning of three frontier large language models (LLMs) across standardized pediatric gastroenterology, hepatology, and nutrition (PGHN) clinical vignettes. METHODS In this cross-sectional study, 25 fictional PGHN vignettes were developed by one board-certified pediatric gastroenterologist and evaluated using three LLMs: Gemini 3.1 Pro, ChatGPT 5.4 Thinking, and Claude Sonnet 4.6 Extended. A conditional two-step prompting protocol was applied. Three blinded, PGHN-certified co-authors independently scored responses using a structured instrument covering four domains (diagnostic accuracy, management, patient safety, reference quality; 0-2 each) and a global quality of clinical reasoning (QCR) score (1-5 Likert scale). Interobserver reliability was assessed using the intraclass correlation coefficient (ICC). Group differences were analyzed using the Kruskal-Wallis H test with Dunn post-hoc correction. RESULTS Interobserver reliability was excellent (ICC: 0.98 for domains; 0.87 for QCR). All three models achieved perfect diagnostic accuracy (median 2.00, interquartile range: 2.00-2.00). Statistically significant intermodel differences were identified in reference quality (p = 0.048) and QCR (p = 0.049). Claude achieved significantly higher QCR scores than ChatGPT (p = 0.043) and demonstrated the highest overall reference quality. Qualitative analysis revealed critical pharmacological dosing errors, contextual blindness, temporal obsolescence, and a high frequency of fabricated citations. CONCLUSION While current LLMs demonstrate good diagnostic pattern recognition in PGHN, reproducible and potentially life-threatening failures in pharmacological reasoning and reference accuracy create a dangerous illusion of competence. These findings suggest that LLMs may support differential diagnosis brainstorming in PGHN but do not establish their safety in real-world clinical scenarios.

Y. Ergen, S. Teke, E. G. Başaran et al. · 0 citations
Jul 2026

Diagnostic and Clinical Management Performance of Large Language Models in Pediatric Infectious Diseases: A Multi-Model Comparative Cross-Sectional Study with Healthcare Professionals and Medical Students

Objective: Large language models (LLMs) are increasingly investigated to support clinical decision-making. However, their comparative performance against human expertise in pediatric infectious diseases remains insufficiently characterized. This study compared diagnostic accuracy and management score between healthcare professionals, medical students, and multiple LLM configurations. Methods: A cross-sectional comparative study was conducted using 30 standardized pediatric infectious disease cases stratified by difficulty (6 easy, 12 medium, 12 difficult). Each case included three multiple-choice questions: a diagnostic gateway question and two management questions. A conditional scoring framework was applied: management questions were scored only when the diagnostic gateway question was answered correctly. The human cohort included 308 participants, generating 2,450 responses. The LLM cohort included 14 model configurations: ChatGPT (5.1 and 5.2, each in Instant and Thinking modes; 5.3 Instant; 5.4 Thinking); Claude (Opus 3, 4.5, 4.6); Gemini (3.1 Pro, Thinking, Fast); and DeepSeek v3 (Normal, DeepThink). Each was evaluated across five independent sessions (2,100 case responses). Results: The LLM cohort demonstrated significantly higher performance than the human cohort across all endpoints. The adjusted difference in diagnostic accuracy was +25.78 percentage points (95% CI: 22.05–29.52; p < 0.001). The overall score was higher by +36.57 points (95% CI: 32.65–40.50; p < 0.001), and the management score by +22.82 points (95% CI: 19.79–25.85; p < 0.001). Among human participants, pediatric specialists achieved the highest performance (overall score: 79.34); no subgroup reached LLM-level performance. Human performance declined with increasing case difficulty, whereas LLM performance remained relatively stable. Variability among LLM configurations was minimal (range: 98.00%–100.00%). Conclusion: Current-generation LLMs demonstrated superior performance in diagnostic accuracy and management score compared with healthcare professionals and medical students on standardized pediatric infectious disease cases. These findings support their potential role as clinical decision support tools. However, further studies are required to evaluate real-world applicability, safety, and human–AI collaboration in clinical practice.

Sait Ramazan Gülbay, Muhammed Yusuf Ozan Avcı, Muhammed Nezih Koç et al. · 0 citations
Review Jul 2026

Guideline Concordance of Large Language Model Responses to Parent-oriented Guideline Prompts About Pediatric Acute Bacterial Arthritis.

BACKGROUND Large language models (LLMs) are increasingly used by patients and caregivers to obtain medical information. In pediatric acute bacterial arthritis, non-guideline-concordant information may be clinically important because timely diagnosis and management are essential. This study evaluated the concordance of LLM-generated responses to parent-oriented reformulations of Pediatric Infectious Diseases Society/Infectious Diseases Society of America (PIDS/IDSA) guideline recommendations. METHODS In this exploratory cross-sectional benchmarking study, 27 PIDS/IDSA guideline-derived recommendations and good practice statements were reformulated into standardized parent-oriented prompts. The same prompts were submitted to GPT-5.4 Thinking, Gemini 3 Thinking and Claude 4.6 Sonnet through browser-based interfaces on April 12, 2026. Responses were anonymized and independently assessed by 3 blinded reviewers. Each response was classified as concordant or discordant; final classifications were determined by majority decision. Interrater agreement was assessed using Fleiss' kappa, and model differences were evaluated using Cochran's Q test. RESULTS Overall, 75 of 81 responses (92.6%) were concordant with PIDS/IDSA recommendations. Gemini 3 Thinking achieved concordance in 27/27 responses (100.0%), Claude 4.6 Sonnet in 25/27 (92.6%) and GPT-5.4 Thinking in 23/27 (85.2%). Cochran's Q test showed no significant difference among models (Q = 4.800, df = 2, P = 0.091). No unsupported or hallucinated content was identified. Interrater agreement was moderate (κ = 0.580; 95% CI, 0.454-0.706; P < 0.001). CONCLUSIONS LLMs showed high guideline concordance, but selected item-level discordance persisted. Because these models may have been trained on guideline-derived content, high concordance should not be equated with independent clinical reasoning. These tools may support caregiver-oriented education but should not replace clinician assessment or guideline-based care.

Ahmet Murat Çörekci, Belen Ateş, Orkun Dinç et al. · 0 citations
Open access Jul 2026

Language-dependent performance variation in large language models for dental trauma management: a comparative evaluation of ChatGPT-5.2, Gemini 3.0, and Claude 4.5 Sonnet.

BACKGROUND Large language models (LLMs) are increasingly evaluated for medical question answering and clinical information tasks, yet the impact of query language on their performance in specialized domains such as dental traumatology remains insufficiently studied. The primary objective was to evaluate whether query language (English vs. Turkish) affects LLM performance in a controlled scenario-based assessment of dental trauma management. Secondary objectives were to compare overall performance across three LLMs and to examine whether language effects are uniform across models or model-specific. METHODS Twenty-seven clinical scenarios covering 13 dental trauma categories were presented to ChatGPT 5.2, Gemini 3.0, and Claude 4.5 Sonnet in both English and Turkish, generating 162 responses. Two blinded endodontists independently evaluated responses using a standardized rubric assessing accuracy (40%), completeness (35%), and safety (25%) against IADT 2020 Guidelines. Inter-rater reliability was assessed using intraclass correlation coefficient (ICC). Language effects were analyzed using Wilcoxon signed-rank tests; model comparisons employed Kruskal-Wallis and Mann-Whitney U tests with Bonferroni correction. RESULTS Inter-rater reliability ranged from moderate to good across evaluation dimensions (ICC: 0.738-0.836). ChatGPT showed the strongest language effect with 9.14% higher performance in English (p < 0.001, r = 0.874). Gemini showed moderate English advantage (5.69%, p = 0.003, r = 0.572). Claude exhibited language independence with virtually identical performance in both languages (-0.02%, p = 0.220). In English, significant model differences emerged (H = 22.31, p < 0.001); however, model performance converged in Turkish (H = 2.89, p = 0.236). CONCLUSIONS This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management. ChatGPT 5.2 achieved the highest performance in English but exhibited the most pronounced Turkish-language degradation, including substantial safety score decline. Gemini 3.0 showed an intermediate pattern with moderate English advantage. Claude 4.5 Sonnet demonstrated language-independent performance across all evaluated dimensions. These findings are based on a standardized scenario-based assessment and should not be extrapolated to real clinical environments or patient care settings.

H. Öz, M. Dundar · 0 citations
Review Aug 2026

Large language model treatment-pathway outputs based on structured clinical text in mid- and low rectal cancer: Concordance with multidisciplinary team decisions and features associated with discordance.

INTRODUCTION This study evaluated concordance between treatment-pathway outputs generated from structured clinical text by the large language model (LLM) GPT-5.4 Thinking and multidisciplinary team (MDT) decisions for mid- and low rectal cancer. MATERIALS AND METHODS This single-center retrospective study included 260 patients who underwent standardized assessment, MDT discussion, and curative-intent surgery between January 2022 and December 2023. After database lock, structured de-identified clinical text was entered into GPT-5.4 Thinking using prespecified preoperative and postoperative templates, with MDT decisions treated as real-world reference decisions. The primary endpoint was preoperative concordance; secondary endpoints included postoperative and overall concordance. Concordance metrics, directional discordance, baseline comparators, repeatability in a 50-case subset, and exploratory logistic regression models were assessed. RESULTS Preoperative, postoperative, and overall concordance rates were 57.7%, 66.9%, and 46.9%, respectively; Cohen's κ values were 0.269 and 0.449 for the preoperative and postoperative stages. Preoperatively, directional discordance relative to MDT decisions included 48 potential under-intensification, 28 potential over-intensification, and 34 directionally indeterminate or heterogeneous cases. Compared with the majority-class baseline, the LLM had the same preoperative crude concordance but higher κ and balanced category-specific concordance; postoperatively, it exceeded the majority-class baseline across these metrics. Non-identical mapped categories across three runs occurred in 12/50 preoperative and 8/50 postoperative assessments despite identical inputs. CONCLUSION GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited. Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs.

Yuanze Wei, Yulong Tian, Xiaodong Liu et al. · 0 citations