Skip to content

Large language models demonstrate variable diagnostic performance and a systematic risk of undertriage in surgical triage of feline metacarpal and metatarsal fractures.

Aug 2026 · Journal of the American Veterinary Medical Association · pp. 1-6 · 0 citations · 20 references
Medicine

Abstract

Objective To evaluate the diagnostic performance of multiple large language models (LLMs) against expert consensus in determining surgical intervention needs for feline metacarpal and metatarsal fractures. Methods In this retrospective study (December 2023 to February 2025), 73 clinical cases of feline metacarpal and metatarsal fractures were evaluated. Two board-certified veterinary orthopedic surgeons established a reference standard for surgical versus conservative management. Five LLMs (ChatGPT, version 5.2 [OpenAI Inc]; Gemini, version 3 Pro [Alphabet Inc]; Grok, version 4.1 [SpaceXAI]; Qwen, version 3.5 [Alibaba Cloud]; and Claude Sonnet, version 4.5 [Anthropic PBC]; and Claude Sonnet, version 4.5 [Anthropic PBC]) assessed anonymized case summaries using a standardized zero-shot prompt. Model recommendations were compared with the reference standard to calculate accuracy, sensitivity, specificity, and the Cohen κ. Results The reference standard classified 49 cases (67.1%) as surgical and 24 (32.9%) as conservative. ChatGPT achieved the highest performance (accuracy, 84.9%; sensitivity, 79.6%; specificity, 95.8%; κ = 0.69). Other models showed lower performance; Qwen and Claude Sonnet failed to identify any surgical cases (0% sensitivity). A systematic bias toward conservative management was observed across all models, causing high false-negative rates (undertriage). Conclusions LLMs demonstrate highly variable diagnostic performance and a systematic risk of undertriage in feline fracture assessment. While top-performing models approach expert-level agreement, others fail in critical clinical scenarios. Clinical Relevance Independent LLM use for surgical decision-making in feline orthopedics is not recommended due to undertriage risks. These tools require strict clinician oversight for preliminary triage.

View source

Similar papers

Open access Jul 2026

Artificial intelligence meets pediatric orthopedics: A comparative analysis of ChatGPT-4o, Gemini 2.0, and Claude 3.5 in detecting supracondylar humeral fractures

Background Supracondylar humeral fractures constitute 10–16% of pediatric skeletal injuries, requiring timely diagnosis to prevent neurovascular complications. Developmental variations in pediatric bone structures pose diagnostic challenges for clinicians. This study evaluated three next-generation large language models (LLMs) (ChatGPT-4o, Gemini 2.0, Claude 3.5) for detecting pediatric supracondylar humeral fractures and their classification according to the Gartland system. Methods This retrospective observational study included 300 pediatric patients (150 with supracondylar humeral fractures confirmed by expert consensus, 150 without fractures) aged 2–10 years presenting to the Emergency Department of the Bilkent City Hospital (October 2022-January 2025). Two-view elbow radiographs were presented to each LLM three times on different days. Diagnostic accuracy was evaluated using overall accuracy (all three responses correct), strict accuracy (≥2 correct responses), and ideal accuracy (≥1 correct response). Response consistency was assessed using Fleiss’ Kappa coefficient. Fractures were classified according to modified Gartland criteria. Results Gemini 2.0 demonstrated highest sensitivity (68.4%) followed by Claude 3.5 (58.7%) and ChatGPT-4o (19.3%) for fracture detection (p < 0.001). Ideal accuracy rates were 83.3%, 78.7%, and 27.3% respectively. Although ideal accuracy rates exceeded 91% in non-fracture cases, specificity remained low (33.1–36.0%), indicating a high rate of false-positive classifications. Response consistency was very good for ChatGPT-4o (κ = 0.69) and Gemini 2.0 (κ = 0.61), good for Claude 3.5 (κ = 0.44). For Gartland classification, Gemini 2.0 achieved highest accuracy: Type I (83.3%), Type II (62.4%), Type III (68.7%). Conclusion Current LLMs demonstrate limited capability as independent diagnostic tools for pediatric supracondylar humeral fractures. Gemini 2.0’s 68.4% sensitivity indicates these technologies require specialized pediatric training before clinical implementation. However, their potential as assistive tools for triage and assessment warrants further development of pediatric-specific models.

U. Kalafat, H. Mutlu, Ramiz Yazıcı et al. · 0 citations
Open access Jul 2026

Language-dependent performance variation in large language models for dental trauma management: a comparative evaluation of ChatGPT-5.2, Gemini 3.0, and Claude 4.5 Sonnet.

BACKGROUND Large language models (LLMs) are increasingly evaluated for medical question answering and clinical information tasks, yet the impact of query language on their performance in specialized domains such as dental traumatology remains insufficiently studied. The primary objective was to evaluate whether query language (English vs. Turkish) affects LLM performance in a controlled scenario-based assessment of dental trauma management. Secondary objectives were to compare overall performance across three LLMs and to examine whether language effects are uniform across models or model-specific. METHODS Twenty-seven clinical scenarios covering 13 dental trauma categories were presented to ChatGPT 5.2, Gemini 3.0, and Claude 4.5 Sonnet in both English and Turkish, generating 162 responses. Two blinded endodontists independently evaluated responses using a standardized rubric assessing accuracy (40%), completeness (35%), and safety (25%) against IADT 2020 Guidelines. Inter-rater reliability was assessed using intraclass correlation coefficient (ICC). Language effects were analyzed using Wilcoxon signed-rank tests; model comparisons employed Kruskal-Wallis and Mann-Whitney U tests with Bonferroni correction. RESULTS Inter-rater reliability ranged from moderate to good across evaluation dimensions (ICC: 0.738-0.836). ChatGPT showed the strongest language effect with 9.14% higher performance in English (p < 0.001, r = 0.874). Gemini showed moderate English advantage (5.69%, p = 0.003, r = 0.572). Claude exhibited language independence with virtually identical performance in both languages (-0.02%, p = 0.220). In English, significant model differences emerged (H = 22.31, p < 0.001); however, model performance converged in Turkish (H = 2.89, p = 0.236). CONCLUSIONS This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management. ChatGPT 5.2 achieved the highest performance in English but exhibited the most pronounced Turkish-language degradation, including substantial safety score decline. Gemini 3.0 showed an intermediate pattern with moderate English advantage. Claude 4.5 Sonnet demonstrated language-independent performance across all evaluated dimensions. These findings are based on a standardized scenario-based assessment and should not be extrapolated to real clinical environments or patient care settings.

H. Öz, M. Dundar · 0 citations
Open access Aug 2026

Diagnostic performance and short-interval reproducibility of multimodal large language models in differentiating cholesteatoma from chronic otitis media using key-image temporal bone high-resolution computed tomography.

PURPOSE This study aimed to assess whether multimodal large language models (LLMs) can distinguish cholesteatoma from non-cholesteatomatous chronic otitis media (COM) on representative key-image temporal bone high-resolution computed tomography (HRCT) and to evaluate the short-interval reproducibility of their outputs. METHODS This retrospective, single-center study (2019-2024) included 101 patients (48 with cholesteatoma, 53 with non-cholesteatomatous COM) who underwent surgical treatment. The reference standard was intraoperative diagnosis with histopathological confirmation for cholesteatoma and surgical documentation for COM. For each case, six anonymized representative HRCT images reflecting standard diagnostic criteria were selected by consensus between a 4th-year radiology resident and a board-certified head and neck radiologist, both unaware of the diagnosis. Subsequently, the same images were analyzed by the GPT-5 and Gemini 2.5 Pro LLMs through their official web interfaces, utilizing structured prompts and a zero-shot approach. These evaluations were conducted in two distinct sessions (S1 and S2) with a 1-week interval. The primary endpoint was accurate binary classification. Accuracy, sensitivity, specificity, positive predictive value, and negative predictive value were calculated with 95% confidence intervals [(CIs); Wilson method] agreement with the reference standard and between sessions and models was assessed with the Cohen kappa (κ) coefficient; and differences in classification were assessed with the McNemar test. RESULTS The radiologist achieved an accuracy of 96.0% (95% CI: 90.3-98.4) with almost perfect agreement with the reference standard (κ: 0.921). In S1, GPT-5 and Gemini 2.5 Pro achieved accuracies of 43.6% and 49.5%, and in S2, 46.5% and 47.5%, respectively. Both models combined high sensitivity (83.3%-97.9%) with low specificity (1.9%-13.2%), and balanced accuracy ranged from 0.45 to 0.52. Between-session reproducibility was fair for GPT-5 (κ: 0.360) and moderate for Gemini 2.5 Pro (κ: 0.485), and inter-model agreement was slight at both sessions (κ: 0.035 at S1 and κ: 0.086 at S2). Accuracy did not differ significantly between the two models (P = 0.211). CONCLUSION In this single-center study, GPT-5 and Gemini 2.5 Pro, in the versions evaluated, combined high sensitivity with low specificity and showed only fair-to-moderate between-session reproducibility and exhibited slight inter-model agreement on temporal bone key-image HRCT. These findings do not support their use as independent second readers, and broader generalization to other multimodal LLMs would require the evaluation of additional models. CLINICAL SIGNIFICANCE The evaluated models lacked the spatial precision and consistency required for the accurate assessment of complex middle ear structures. This finding underscores the necessity for verification by a radiologist and continuous monitoring.

B. Yağcı, Sergen Palaz, E.A. Cetinkaya et al. · 0 citations
Review Open access Aug 2026

Artificial Intelligence for Diagnosis of Temporomandibular and Cranio-Cervico-Mandibular Musculoskeletal Disorders: A Systematic Review and Exploratory Diagnostic Test Accuracy Meta-Analysis

Objectives: To systematically evaluate the diagnostic accuracy, clinical applicability, and methodological maturity of artificial intelligence (AI)-based methods for temporomandibular disorders (TMD), temporomandibular joint (TMJ) abnormalities, and related cranio-cervico-mandibular (CCM) musculoskeletal conditions compared with conventional diagnostic methods and accepted reference standards. Materials and Methods: This systematic review and exploratory diagnostic test accuracy meta-analysis was conducted in accordance with PRISMA 2020 and PRISMA-DTA. The protocol was retrospectively registered in PROSPERO (CRD420261428138). PubMed/MEDLINE, Embase, and Scopus were searched from database inception through February 2026. Eligibility for the primary synthesis was restricted to published studies in English or Spanish involving adults aged 18 years or older. All extracted records were re-audited article by article to align the evidence with the diagnostic question. The domain-specific quantitative synthesis was restricted to TMJ osteoarthritis studies with explicit 2 × 2 diagnostic data or a unique, verifiable reconstruction from reported class totals and sensitivity/specificity. Risk of bias was assessed with QUADAS-2. Results: From 1471 records identified, 174 entered the master extraction dataset. After reclassification, 84 records were retained for primary TMD/TMJ qualitative synthesis, 8 as secondary CCM musculoskeletal evidence, 31 as conventional or reference standard supporting evidence, 33 as methodological or contextual evidence, 4 as differential orofacial pain evidence, and 14 as excluded or minimal-background records. Twenty-one studies were assessed as potential diagnostic accuracy candidates. Three TMJ osteoarthritis studies contributed to the domain-specific exploratory meta-analysis: two with explicit 2 × 2 data and one with a reproducible reconstruction. Pooled sensitivity was 0.791 (95% CI: 0.700–0.861) and pooled specificity was 0.869 (95% CI: 0.811–0.911). Heterogeneity was substantial for sensitivity (I2 = 68.2%) and moderate for specificity (I2 = 57.3%). Conclusions: AI demonstrates promising performance in selected image-based TMJ osteoarthritis tasks. Nevertheless, the evidence remains exploratory because only three studies were quantitatively comparable, one table was reconstructed, and modalities and validation designs differed. AI should be interpreted as an augmentative decision support tool rather than a replacement for MRI, CBCT, or validated clinical frameworks such as DC/TMD. Clinical Relevance: AI may support image-based TMD/TMJ workflows, but present evidence does not justify autonomous diagnosis or replacement of established clinical and imaging reference standards.

Arturo Arbeláez Ramírez, Daniel Botero Rosas · 0 citations