Skip to content

Diagnostic accuracy of large language models in ICOP-based orofacial pain diagnosis: A comparative study.

Jul 2026 · Cranio · pp. 1-9 · 0 citations · 21 references
Medicine

TL;DR

Although Grok 4 showed the highest diagnostic concordance, current LLMs should support, not replace, clinician judgment, and LLM performance varied across ICOP-based scenarios.

Abstract

Objective

To compare the diagnostic performance of ChatGPT 5.5, Claude Opus 4.1, Gemini 3 Flash, and Grok 4 in International Classification of Orofacial Pain (ICOP)-based clinical scenarios.

Methods

Thirty ICOP diagnoses were randomly selected, and corresponding clinical scenarios were manually developed. Each scenario was submitted to all models using standardized prompts in independent sessions. Two blinded evaluators assessed primary diagnosis accuracy, subclassification accuracy, clinical interpretation, and management recommendations.

Results

Overall performance differed significantly among models (p < .001). Grok 4 achieved the highest total score and outperformed the other models. No significant differences were found among ChatGPT 5.5, Gemini 3 Flash, and Claude Opus 4.1. Subclassification accuracy was consistently lower than primary diagnosis accuracy, while management recommendations did not differ significantly.

Conclusion

LLM performance varied across ICOP-based scenarios. Although Grok 4 showed the highest diagnostic concordance, current LLMs should support, not replace, clinician judgment.

View source

Similar papers

Open access Aug 2026

Accuracy and Completeness of Contemporary Large Language Models in Prosthodontics: An Expert-Based Comparative Study

Purpose: To compare the scientific accuracy and response completeness of four contemporary LLMs: ChatGPT-5.2, Gemini 3, Copilot, and DeepSeek-V3.2 across major prosthodontic domains and question formats.Methods: Fifty prosthodontic questions covering removable prosthodontics, fixed prosthodontics, implantology, temporomandibular disorders and occlusion, dental materials science were developed from standard textbooks and clinical guidelines. Questions were presented in yes/no, multiple-choice, and open-ended formats to reflect different levels of clinical reasoning. Responses generated by ChatGPT-5.2 (OpenAI), Gemini 3 (Google), Microsoft Copilot (Microsoft), and DeepSeek-V3.2 were independently evaluated by five prosthodontists using Likert-based scales for accuracy and completeness. Inter-rater reliability was assessed to ensure scoring consistency, and model performances were compared across domains and question types using nonparametric statistical tests.Results: Overall accuracy differed significantly among the four models (p0.05), although descriptive trends suggested slightly higher scores for yes/no questions and lower scores for multiple-choice items. Response completeness also varied significantly across models (p0.05).Conclusion: Contemporary large language models demonstrate acceptable theoretical performance in prosthodontics; however, clinically relevant differences persist in response completeness and consistency. These models may support prosthodontic education and preliminary information retrieval, but cannot replace expert clinical judgment.

Elif Yiğit İren, Hatice Betül Üçkuyu · 0 citations
Open access Aug 2026

Accuracy and Consistency of Three Large Language Models on Fixed Prosthodontics Questions

Large language models (LLMs) are increasingly used in dental education and clinical settings, but their accuracy and consistency in fixed prosthodontics remain uncertain. Objectives: To compare the accuracy (using a strict three-attempt criterion) and repeated response consistency of ChatGPT-5.4, Gemini 3.1 Pro, and Claude Opus 4.8 in answering binary fixed prosthodontics questions across six clinical domains. Methods: Forty binary questions covering six core fixed prosthodontic domains were developed by the principal investigator. Two experienced faculty members independently answered the questions, and disagreements were resolved through discussion to establish the gold standard. Each question was submitted thrice to each LLM in a new chat session on the same day. Accuracy was the percentage of correct responses compared with the gold standard across all three attempts, while consistency was the percentage of identical responses across three attempts, regardless of correctness. Accuracy and consistency were expressed as percentages with 95% confidence intervals. Cochran’s Q test, Cohen’s kappa, and Fleiss’ kappa were used for statistical analysis. Results: Pre-consensus examiner agreement was moderate, 77.5% (Cohen’s κ=0.55). Claude Opus 4.8 achieved the highest accuracy (95% and consistency (97.5%, followed by ChatGPT-5.4 87.5% accuracy and 92.5% consistency and Gemini 3.1 Pro (85% accuracy and 87.5% consistency. No significant difference was observed in accuracy among models (Q=5.20, p=0.074). Intra-model consistency for three models was almost perfect (κ=0.83–0.97). Conclusions: The evaluated LLMs demonstrated high accuracy and consistency for structured fixed prosthodontics questions, with Claude Opus 4.8 highest, though not significant.

A. Bashir, Moeen Ud Din Ahmad, Ussamah Waheed Jatala et al. · 0 citations
Open access Jul 2026

Comparison of responses from large language models using artificial intelligence for parent-focused inquiries on clubfoot treatment and Ponseti management

This study aimed to compare the accuracy, completeness, and readability of responses generated by three large language models (LLMs)—ChatGPT-4.0 (OpenAI), Microsoft CoPilot, and Google Gemini—regarding the treatment and management of idiopathic congenital talipes equinovarus (ICTEV) using the Ponseti method. Fifteen frequently asked questions were selected from pediatric orthopedic center websites, Google Trends analysis, and clinical experience. Each question was submitted verbatim in a new session to the three LLMs within 24 h. Seven board-certified pediatric orthopedic surgeons, blinded to the source, rated responses for accuracy (5-point Likert scale) and completeness (3-point Likert scale). Readability was assessed using the Flesch–Kincaid grade level. Mean scores ± standard deviation were calculated, and inter-rater reliability was estimated using the intraclass correlation coefficient (ICC). Group differences were tested with ANOVA and chi-squared tests ( p  < 0.05). A total of 45 responses were evaluated. Gemini achieved the highest mean accuracy (4.1 ± 0.8), followed by CoPilot (3.6 ± 0.8) and ChatGPT-4.0 (3.3 ± 0.9), with significant differences among models ( p  < 0.001). Completeness ratings also differed significantly ( p  < 0.001). Readability analysis showed that ChatGPT produced shorter, more readable text, while Gemini generated longer, more complex responses. Inter-rater reliability was substantial for accuracy (ICC 0.715) and completeness (ICC 0.710). Google Gemini outperformed ChatGPT-4.0 and CoPilot in accuracy and comprehensiveness for ICTEV management using the Ponseti method. However, its complex responses may limit accessibility, whereas ChatGPT-4.0 offered more readable but less detailed answers. IV

A. Vescio, G. Testa, M. Sapienza et al. · 0 citations
Review Open access Jul 2026

Evaluating large language models as tools for public health education on scoliosis

Purpose This study was aimed to compare the efficacy of three most popular large language models (LLMs)—Claude Opus 4.6, ChatGPT Thinking 5.4 and DeepSeek v3.2 in answering frequently asked questions (FAQs) about scoliosis. Methods 20 scoliosis related questions (four categories, five questions in each category) were submitted to each LLM. A panel of 9 experts (two spine surgeons, two pediatric orthopedic surgeons and five physical therapists, all blinded to the LLMs and responses) rated independently each response generated by LLMs on a 6 points Likert scale (1 as strongly disagree to 6 as strongly agree). 540 total ratings were collected. Intergroup comparisons were conducted by Kruskal Wallis test and Mann Whitney U pairwise tests. Paired question level analysis was achieved by Friedman test and Wilcoxon signed rank comparisons. Results Claude’s score was 5.53 ± 0.76 much higher than both ChatGPT (4.84 ± 0.86, p < 0.001) and DeepSeek (4.86 ± 0.84, p < 0.001), but no difference was found between ChatGPT and DeepSeek (p = 0.749). Claude performed on top for 19 of 20 questions (95%) and was favored by 7 of 9 reviewers. Consistency of Claude was also highest [CV = 13.8% vs. 17.9% (ChatGPT) and 17.2% (DeepSeek)]. Conclusion Although all three LLMs achieved favorable overall ratings (>4.8/6), Claude performed significantly better than ChatGPT and DeepSeek for scoliosis FAQs taking into account its higher accuracy and consistency. Within the scope of the present evaluation, Claude demonstrated the strongest overall performance among the three LLMs tested.

Yi-Chen Wang, Xuanyan Liu, Mengjie Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.