Large Language Models and Retrieval-Augmented Platforms for the Diagnosis and Management of Periodontal Diseases: A Blinded Expert-Rated Comparative Study of 11 Systems.
Jul 2026· Journal of Clinical Periodontology· 0 citations· 19 references
Medicine
TL;DR
Retrieval-augmented systems rated highest, but this advantage was confounded with response length, and the dangerous-response spread argues against undifferentiated use.
Abstract
Aim
To compare retrieval-augmented systems with general-purpose large language models (LLMs) on standardised periodontal clinical vignettes.
Materials And Methods
Eleven AI systems were evaluated: nine general-purpose LLMs, one general-purpose retrieval-augmented platform (Perplexity) and one medical-domain retrieval-augmented platform (OpenEvidence). Each responded to 30 synthetic vignettes covering acute, chronic and complex periodontal scenarios. Six blinded periodontists scored responses on a 5-point Likert scale for accuracy, safety, freedom from hallucinations and completeness in a randomised block design. Friedman and Conover-Iman tests with Holm correction were applied; mixed-effects and ordinal models served as sensitivity analyses.
Results
At least one parameter scored dangerous (≤ 2) in 3.3%-46.7% of responses across platforms, despite mean composite scores (3.28-4.86) exceeding the rubric midpoint of 3.0. Between-model differences were significant (p < 0.001), with a small-to-medium overall effect (Kendall's W = 0.17) and large within-category effects (W up to 0.82). Perplexity, OpenEvidence and Claude 4.7 Opus formed a top tier.
Conclusion
Retrieval-augmented systems rated highest, but this advantage was confounded with response length. The dangerous-response spread argues against undifferentiated use. These tools should assist, not replace, specialist judgement.
INTRODUCTION
Large language models (LLMs) are increasingly utilized for medical and dental information retrieval, yet their ability to interpret authentic, patient-style inquiries remains insufficiently investigated. This study compared the performance of ChatGPT, Claude, and Gemini in responding to patient-oriented queries related to periodontal and peri-implant diseases.
MATERIALS AND METHODS
Unlike traditional investigations using expert-generated questions, this study employed 40 realistic, patient-oriented queries designed to simulate the post-examination cognitive state, blending colloquial language with partially retained clinical jargon. Each query was submitted to GPT-4o, Claude Sonnet 5 and Gemini 2.5 Pro generating 120 responses. Three blinded periodontists independently evaluated scientific accuracy, completeness, clinical safety, and overall quality using a 5-point Likert scale. Automated text analysis assessed readability metrics (Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index) and linguistic characteristics. Statistical protocols included Friedman, Bonferroni-adjusted Wilcoxon signed-rank, and Model Dominance analyses.
RESULTS
Significant performance differences were observed among the models across all expert-rated domains (all p < 0.001). Gemini achieved the highest expert ratings for scientific accuracy (4.77 ± 0.22), clinical safety (4.94 ± 0.15), and overall quality (4.85 ± 0.20), and was identified as the most frequently top-ranked platform via dominance analysis. Conversely, Claude performed significantly better regarding response completeness (4.77 ± 0.22) and demonstrated the most favorable overall readability profile, yielding the lowest Flesch-Kincaid Grade Level (7.54 ± 1.27). GPT-4o consistently received the lowest expert ratings across all evaluated domains.
DISCUSSION
While all evaluated LLMs generated high-quality responses to realistic periodontal queries, their functional strengths were highly multidimensional. Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations. These findings support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical oversight.
Ramazan Ağırağaç, Vedat Yüksekkaya· Journal of Stomatology Oral...· 0 citations
Purpose: To compare the scientific accuracy and response completeness of four contemporary LLMs: ChatGPT-5.2, Gemini 3, Copilot, and DeepSeek-V3.2 across major prosthodontic domains and question formats.Methods: Fifty prosthodontic questions covering removable prosthodontics, fixed prosthodontics, implantology, temporomandibular disorders and occlusion, dental materials science were developed from standard textbooks and clinical guidelines. Questions were presented in yes/no, multiple-choice, and open-ended formats to reflect different levels of clinical reasoning. Responses generated by ChatGPT-5.2 (OpenAI), Gemini 3 (Google), Microsoft Copilot (Microsoft), and DeepSeek-V3.2 were independently evaluated by five prosthodontists using Likert-based scales for accuracy and completeness. Inter-rater reliability was assessed to ensure scoring consistency, and model performances were compared across domains and question types using nonparametric statistical tests.Results: Overall accuracy differed significantly among the four models (p0.05), although descriptive trends suggested slightly higher scores for yes/no questions and lower scores for multiple-choice items. Response completeness also varied significantly across models (p0.05).Conclusion: Contemporary large language models demonstrate acceptable theoretical performance in prosthodontics; however, clinically relevant differences persist in response completeness and consistency. These models may support prosthodontic education and preliminary information retrieval, but cannot replace expert clinical judgment.
This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management.
INTRODUCTION
The use of artificial intelligence (AI) in orthodontic practice is increasing rapidly; however, there is a notable lack of research evaluating the accuracy of large language models (LLMs) in educating patients about orthodontic retainers and related guidelines.
MATERIALS AND METHODS
This study utilized a cross-sectional, repeated‑measures comparative evaluation design after receiving exemption from the institutional ethics committee. A set of 110 questions related to orthodontic retainers was compiled from previous articles addressing concerns about retainers and approved by a panel of three orthodontists. These questions were submitted to large language models (LLMs), including ChatGPT, Copilot, DeepSeek and Google Gemini. The responses were then reviewed by six independent orthodontists, who rated them using a modified five-point Likert scale.
RESULTS
The overall accuracy revealed that 68.6% of responses scored 4, while 15.3% achieved a perfect score of 5. Among the LLMs, Gemini ranked first with 96.8%, closely followed by ChatGPT at 95.6%, indicating comparable high‑level performance between these models, while DeepSeek (76.9%) and Copilot (66.2%) demonstrated comparatively lower accuracy. Gemini produced a higher proportion of perfect scores, whereas ChatGPT consistently achieved strong ratings. The mean ratings across six raters demonstrated strong reliability (ICC = 0.81), reflecting expert agreement.
CONCLUSIONS
Findings suggest that AI models such as ChatGPT and Gemini can generate patient‑directed orthodontic retainer information with high informational accuracy under controlled evaluation conditions. However, specialist oversight remains essential to ensure clinical applicability. Future research using larger and more diverse datasets is needed to assess broader educational and communication‑related outcomes.
Muhammad Mughni, M. Ilyas, Syed Ali Naqi Gilani et al.· BMC Oral Health· 0 citations
BACKGROUND
Large language models (LLMs) are increasingly used by clinicians and learners for endodontic information, yet their reliability for irrigation-related knowledge remains unclear. This study evaluated the accuracy, readability, hallucination profile, clinical risk, and temporal stability of 4 general-purpose LLMs on endodontic irrigation questions, including false-premise prompts.
METHODS
A literature-based reference set was developed for sodium hypochlorite, calcium hypochlorite, ethylenediaminetetraacetic acid, and chlorhexidine. ChatGPT-5.2, Claude Sonnet 4.5, Gemini 3 Pro, and DeepSeek 3.2 were each asked 100 questions comprising 80 factual items and 20 contradiction-seeking items. Responses were independently scored by 2 blinded endodontists for accuracy, hallucination subtype, and clinical risk. Five readability indices were calculated, and model stability was reassessed 10 days later.
RESULTS
Inter-rater agreement was almost perfect (weighted κ = 0.92 for accuracy; κ = 0.97 for hallucination). Claude Sonnet 4.5 and Gemini 3 Pro showed the highest accuracy (1.75 and 1.74) and the lowest hallucination rates (7% and 9%). DeepSeek 3.2 showed the lowest accuracy (1.08), the highest hallucination rate (29%), and critical-risk outputs in 16% of responses. Gemini showed the highest test-retest stability (weighted κ = 0.95). Hallucination strongly correlated with clinical risk (ρ = 0.93; P < 0.001). Readability analysis showed a 2-tier pattern: Gemini and DeepSeek produced more accessible text, whereas ChatGPT and Claude generated denser outputs.
CONCLUSIONS
LLM performance in endodontic irrigation is model- and irrigant-dependent. Hallucination profiling is clinically relevant, and high initial accuracy does not guarantee temporal stability. LLM outputs should be used only as clinician-verified adjuncts.