Comparative Evaluation of Large Language Models as Virtual Orthodontic Patient Information Assistants
Abstract
Objective: The aim of this study was to evaluate the potential use of widely used large language models (LLMs) in patient education by comparing their responses to orthodontic patient questions in terms of accuracy, comprehensiveness, and readability.Methods: This study evaluated the performance of Claude Sonnet 4.6, ChatGPT 5.5 and Gemini 3.5 Pro in providing orthodontic patient information.Ten frequently asked questions by orthodontic patients were posed to each model in independent sessions using the same wording. The responses obtained were scored by two orthodontists for accuracy and comprehensiveness using a modified Likert scale. The readability of the responses was calculated using the FRE and FKGL indices. Differences between the models were evaluated with statistical analyses appropriate for the data structure (p < 0.05).Results: No significant difference was found between the models regarding readability. However, significant differences were found in accuracy and comprehensiveness scores. Gemini 3.5 Pro achieved the highest scores in terms of comprehensiveness and accuracy showed significant superiority over ChatGPT 5.5 (p = 0.007).Conclusion: This study demonstrates that LLMs can serve as auxiliary tools in providing orthodontic patient information.However, responses should be checked by an expert before being presented to patients and should be adapted into simpler and more understandable language.