Skip to content
Open access

Large Language Models and Retrieval-Augmented Platforms for the Diagnosis and Management of Periodontal Diseases: A Blinded Expert-Rated Comparative Study of 11 Systems.

Jul 2026 · Journal of Clinical Periodontology · 0 citations · 19 references
Medicine

TL;DR

Retrieval-augmented systems rated highest, but this advantage was confounded with response length, and the dangerous-response spread argues against undifferentiated use.

Abstract

Aim

To compare retrieval-augmented systems with general-purpose large language models (LLMs) on standardised periodontal clinical vignettes.

Materials And Methods

Eleven AI systems were evaluated: nine general-purpose LLMs, one general-purpose retrieval-augmented platform (Perplexity) and one medical-domain retrieval-augmented platform (OpenEvidence). Each responded to 30 synthetic vignettes covering acute, chronic and complex periodontal scenarios. Six blinded periodontists scored responses on a 5-point Likert scale for accuracy, safety, freedom from hallucinations and completeness in a randomised block design. Friedman and Conover-Iman tests with Holm correction were applied; mixed-effects and ordinal models served as sensitivity analyses.

Results

At least one parameter scored dangerous (≤ 2) in 3.3%-46.7% of responses across platforms, despite mean composite scores (3.28-4.86) exceeding the rubric midpoint of 3.0. Between-model differences were significant (p < 0.001), with a small-to-medium overall effect (Kendall's W = 0.17) and large within-category effects (W up to 0.82). Perplexity, OpenEvidence and Claude 4.7 Opus formed a top tier.

Conclusion

Retrieval-augmented systems rated highest, but this advantage was confounded with response length. The dangerous-response spread argues against undifferentiated use. These tools should assist, not replace, specialist judgement.

Read PDF

Similar papers

Aug 2026

Comparative Evaluation of Large Language Models in Answering Patient Questions Following Periodontal and Peri-Implant Examination: An Expert-Based Study.

INTRODUCTION Large language models (LLMs) are increasingly utilized for medical and dental information retrieval, yet their ability to interpret authentic, patient-style inquiries remains insufficiently investigated. This study compared the performance of ChatGPT, Claude, and Gemini in responding to patient-oriented queries related to periodontal and peri-implant diseases. MATERIALS AND METHODS Unlike traditional investigations using expert-generated questions, this study employed 40 realistic, patient-oriented queries designed to simulate the post-examination cognitive state, blending colloquial language with partially retained clinical jargon. Each query was submitted to GPT-4o, Claude Sonnet 5 and Gemini 2.5 Pro generating 120 responses. Three blinded periodontists independently evaluated scientific accuracy, completeness, clinical safety, and overall quality using a 5-point Likert scale. Automated text analysis assessed readability metrics (Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index) and linguistic characteristics. Statistical protocols included Friedman, Bonferroni-adjusted Wilcoxon signed-rank, and Model Dominance analyses. RESULTS Significant performance differences were observed among the models across all expert-rated domains (all p < 0.001). Gemini achieved the highest expert ratings for scientific accuracy (4.77 ± 0.22), clinical safety (4.94 ± 0.15), and overall quality (4.85 ± 0.20), and was identified as the most frequently top-ranked platform via dominance analysis. Conversely, Claude performed significantly better regarding response completeness (4.77 ± 0.22) and demonstrated the most favorable overall readability profile, yielding the lowest Flesch-Kincaid Grade Level (7.54 ± 1.27). GPT-4o consistently received the lowest expert ratings across all evaluated domains. DISCUSSION While all evaluated LLMs generated high-quality responses to realistic periodontal queries, their functional strengths were highly multidimensional. Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations. These findings support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical oversight.

Ramazan Ağırağaç, Vedat Yüksekkaya · 0 citations
Open access Aug 2026

Accuracy and Completeness of Contemporary Large Language Models in Prosthodontics: An Expert-Based Comparative Study

Purpose: To compare the scientific accuracy and response completeness of four contemporary LLMs: ChatGPT-5.2, Gemini 3, Copilot, and DeepSeek-V3.2 across major prosthodontic domains and question formats.Methods: Fifty prosthodontic questions covering removable prosthodontics, fixed prosthodontics, implantology, temporomandibular disorders and occlusion, dental materials science were developed from standard textbooks and clinical guidelines. Questions were presented in yes/no, multiple-choice, and open-ended formats to reflect different levels of clinical reasoning. Responses generated by ChatGPT-5.2 (OpenAI), Gemini 3 (Google), Microsoft Copilot (Microsoft), and DeepSeek-V3.2 were independently evaluated by five prosthodontists using Likert-based scales for accuracy and completeness. Inter-rater reliability was assessed to ensure scoring consistency, and model performances were compared across domains and question types using nonparametric statistical tests.Results: Overall accuracy differed significantly among the four models (p0.05), although descriptive trends suggested slightly higher scores for yes/no questions and lower scores for multiple-choice items. Response completeness also varied significantly across models (p0.05).Conclusion: Contemporary large language models demonstrate acceptable theoretical performance in prosthodontics; however, clinically relevant differences persist in response completeness and consistency. These models may support prosthodontic education and preliminary information retrieval, but cannot replace expert clinical judgment.

Elif Yiğit İren, Hatice Betül Üçkuyu · 0 citations
Review Open access Jul 2026

Evaluating the informational accuracy of large language models in patient‑directed orthodontic retainer guidance: a cross‑sectional comparison of Chat GPT‑4.1, Gemini 2.5, Microsoft Copilot GPT-4.1 and DeepSeek‑V3.

INTRODUCTION The use of artificial intelligence (AI) in orthodontic practice is increasing rapidly; however, there is a notable lack of research evaluating the accuracy of large language models (LLMs) in educating patients about orthodontic retainers and related guidelines. MATERIALS AND METHODS This study utilized a cross-sectional, repeated‑measures comparative evaluation design after receiving exemption from the institutional ethics committee. A set of 110 questions related to orthodontic retainers was compiled from previous articles addressing concerns about retainers and approved by a panel of three orthodontists. These questions were submitted to large language models (LLMs), including ChatGPT, Copilot, DeepSeek and Google Gemini. The responses were then reviewed by six independent orthodontists, who rated them using a modified five-point Likert scale. RESULTS The overall accuracy revealed that 68.6% of responses scored 4, while 15.3% achieved a perfect score of 5. Among the LLMs, Gemini ranked first with 96.8%, closely followed by ChatGPT at 95.6%, indicating comparable high‑level performance between these models, while DeepSeek (76.9%) and Copilot (66.2%) demonstrated comparatively lower accuracy. Gemini produced a higher proportion of perfect scores, whereas ChatGPT consistently achieved strong ratings. The mean ratings across six raters demonstrated strong reliability (ICC = 0.81), reflecting expert agreement. CONCLUSIONS Findings suggest that AI models such as ChatGPT and Gemini can generate patient‑directed orthodontic retainer information with high informational accuracy under controlled evaluation conditions. However, specialist oversight remains essential to ensure clinical applicability. Future research using larger and more diverse datasets is needed to assess broader educational and communication‑related outcomes.

Muhammad Mughni, M. Ilyas, Syed Ali Naqi Gilani et al. · 0 citations
Aug 2026

Evaluating Large Language Models in Endodontic Irrigants: A Structured Assessment of Accuracy, Readability, Hallucination Subtypes and Clinical Risk.

BACKGROUND Large language models (LLMs) are increasingly used by clinicians and learners for endodontic information, yet their reliability for irrigation-related knowledge remains unclear. This study evaluated the accuracy, readability, hallucination profile, clinical risk, and temporal stability of 4 general-purpose LLMs on endodontic irrigation questions, including false-premise prompts. METHODS A literature-based reference set was developed for sodium hypochlorite, calcium hypochlorite, ethylenediaminetetraacetic acid, and chlorhexidine. ChatGPT-5.2, Claude Sonnet 4.5, Gemini 3 Pro, and DeepSeek 3.2 were each asked 100 questions comprising 80 factual items and 20 contradiction-seeking items. Responses were independently scored by 2 blinded endodontists for accuracy, hallucination subtype, and clinical risk. Five readability indices were calculated, and model stability was reassessed 10 days later. RESULTS Inter-rater agreement was almost perfect (weighted κ = 0.92 for accuracy; κ = 0.97 for hallucination). Claude Sonnet 4.5 and Gemini 3 Pro showed the highest accuracy (1.75 and 1.74) and the lowest hallucination rates (7% and 9%). DeepSeek 3.2 showed the lowest accuracy (1.08), the highest hallucination rate (29%), and critical-risk outputs in 16% of responses. Gemini showed the highest test-retest stability (weighted κ = 0.95). Hallucination strongly correlated with clinical risk (ρ = 0.93; P < 0.001). Readability analysis showed a 2-tier pattern: Gemini and DeepSeek produced more accessible text, whereas ChatGPT and Claude generated denser outputs. CONCLUSIONS LLM performance in endodontic irrigation is model- and irrigant-dependent. Hallucination profiling is clinically relevant, and high initial accuracy does not guarantee temporal stability. LLM outputs should be used only as clinician-verified adjuncts.

Damla Erkal, Yunus Emre Çakmak, Kürşat Er · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.