Author

A. Öztürk

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Jul 2026

Guideline Concordance of Large Language Model Responses to Parent-oriented Guideline Prompts About Pediatric Acute Bacterial Arthritis.

BACKGROUND Large language models (LLMs) are increasingly used by patients and caregivers to obtain medical information. In pediatric acute bacterial arthritis, non-guideline-concordant information may be clinically important because timely diagnosis and management are essential. This study evaluated the concordance of LLM-generated responses to parent-oriented reformulations of Pediatric Infectious Diseases Society/Infectious Diseases Society of America (PIDS/IDSA) guideline recommendations. METHODS In this exploratory cross-sectional benchmarking study, 27 PIDS/IDSA guideline-derived recommendations and good practice statements were reformulated into standardized parent-oriented prompts. The same prompts were submitted to GPT-5.4 Thinking, Gemini 3 Thinking and Claude 4.6 Sonnet through browser-based interfaces on April 12, 2026. Responses were anonymized and independently assessed by 3 blinded reviewers. Each response was classified as concordant or discordant; final classifications were determined by majority decision. Interrater agreement was assessed using Fleiss' kappa, and model differences were evaluated using Cochran's Q test. RESULTS Overall, 75 of 81 responses (92.6%) were concordant with PIDS/IDSA recommendations. Gemini 3 Thinking achieved concordance in 27/27 responses (100.0%), Claude 4.6 Sonnet in 25/27 (92.6%) and GPT-5.4 Thinking in 23/27 (85.2%). Cochran's Q test showed no significant difference among models (Q = 4.800, df = 2, P = 0.091). No unsupported or hallucinated content was identified. Interrater agreement was moderate (κ = 0.580; 95% CI, 0.454-0.706; P < 0.001). CONCLUSIONS LLMs showed high guideline concordance, but selected item-level discordance persisted. Because these models may have been trained on guideline-derived content, high concordance should not be equated with independent clinical reasoning. These tools may support caregiver-oriented education but should not replace clinician assessment or guideline-based care.

Ahmet Murat Çörekci, Belen Ateş, Orkun Dinç et al. · 0 citations