Aug 2026· Turkish Journal of Orthodontics· 0 citations
Medicine
TL;DR
While XGBoost yielded the highest overall classification accuracy, ChatGPT-5.0's high sensitivity in detecting extraction cases was noteworthy, and it demonstrated performance comparable to, and in some cases superior to, traditional ML models for orthodontic extraction decisions.
Abstract
Objective
To make accurate orthodontic extraction decisions, various clinical and cephalometric variables must be evaluated. This study aims to evaluate ChatGPT-5.0's performance in distinguishing orthodontic extraction decisions and to compare it with five supervised machine learning (ML) algorithms.
Methods
Of 550 retrospectively evaluated orthodontic records, 30 were reserved for calibration, leaving 520 for the main analysis. The reference standard was the consensus treatment decision of three expert orthodontists with more than 5 years of clinical experience. Overall, 23 variables were analyzed, including 13 clinical parameters, 7 cephalometric measurements, and photographs. ChatGPT-5.0's performance was evaluated using a 5-fold cross-validation design. It was compared with XGBoost, random forest, support vector machine (SVM), logistic regression, and multi-layer perceptron (MLP). Performance metrics included accuracy, sensitivity, specificity, precision, F1-score, and balanced accuracy, with 95% confidence intervals calculated. Statistical analyses utilized Cochran's Q test and the McNemar test with Holm-Bonferroni correction.
Results
Of the 520 main cases, 223 (42.88%) were extraction treatments and 297 (57.12%) were non-extraction treatments. XGBoost achieved the highest accuracy (78.08%), followed closely by ChatGPT-5.0 (75.77%). The overall performance difference among models was significant (p≤0.001). In pairwise comparisons, ChatGPT's accuracy was significantly higher than those of random forest, SVM, logistic regression, and MLP, but was found to be similar to XGBoost. ChatGPT-5.0 showed the highest sensitivity (76.68%), whereas XGBoost showed the highest specificity (82.15%).
Conclusion
ChatGPT-5.0 demonstrated performance comparable to, and in some cases superior to, traditional ML models for orthodontic extraction decisions. While XGBoost yielded the highest overall classification accuracy, ChatGPT-5.0's high sensitivity in detecting extraction cases was noteworthy.
INTRODUCTION
The use of artificial intelligence (AI) in orthodontic practice is increasing rapidly; however, there is a notable lack of research evaluating the accuracy of large language models (LLMs) in educating patients about orthodontic retainers and related guidelines.
MATERIALS AND METHODS
This study utilized a cross-sectional, repeated‑measures comparative evaluation design after receiving exemption from the institutional ethics committee. A set of 110 questions related to orthodontic retainers was compiled from previous articles addressing concerns about retainers and approved by a panel of three orthodontists. These questions were submitted to large language models (LLMs), including ChatGPT, Copilot, DeepSeek and Google Gemini. The responses were then reviewed by six independent orthodontists, who rated them using a modified five-point Likert scale.
RESULTS
The overall accuracy revealed that 68.6% of responses scored 4, while 15.3% achieved a perfect score of 5. Among the LLMs, Gemini ranked first with 96.8%, closely followed by ChatGPT at 95.6%, indicating comparable high‑level performance between these models, while DeepSeek (76.9%) and Copilot (66.2%) demonstrated comparatively lower accuracy. Gemini produced a higher proportion of perfect scores, whereas ChatGPT consistently achieved strong ratings. The mean ratings across six raters demonstrated strong reliability (ICC = 0.81), reflecting expert agreement.
CONCLUSIONS
Findings suggest that AI models such as ChatGPT and Gemini can generate patient‑directed orthodontic retainer information with high informational accuracy under controlled evaluation conditions. However, specialist oversight remains essential to ensure clinical applicability. Future research using larger and more diverse datasets is needed to assess broader educational and communication‑related outcomes.
Muhammad Mughni, M. Ilyas, Syed Ali Naqi Gilani et al.· BMC Oral Health· 0 citations
INTRODUCTION
To evaluate the reliability and between-method agreement of a freely available orthodontic model analysis software (FusionAnalyser) against a commercial platform (OrthoAnalyzer™) and conventional digital calliper measurements on plaster casts.
MATERIAL AND METHODS
In this retrospective cross-sectional study, pre-treatment records of 40 patients collected between 2022 and 2025 were analysed by two examiners, each performing measurements twice at least 15 days apart. Mesiodistal tooth widths, Bolton ratios, and arch space deficiency were measured by all three methods. Intra- and inter-rater reliability were assessed using intraclass correlation coefficients and Dahlberg errors; between-method agreement was assessed with Friedman and Wilcoxon signed-rank tests, and Bland-Altman analyses with predefined clinical thresholds (±0.50mm for tooth widths, ±1.50 percentage points for Bolton ratios, ±1.00mm for arch space deficiency).
RESULTS
Of 52 records screened, 40 met the eligibility criteria and were included. ICC values across all methods and both examiners ranged from 0.74 (95% CI 0.56-0.86) to 0.996 (95% CI 0.993-0.998). Between-method differences in mesiodistal tooth widths were small (mostly 0.10-0.30mm) and remained within the ±0.50mm threshold. Limits of agreement for the anterior Bolton ratio exceeded ±1.50 percentage points across all comparisons. For mandibular arch space deficiency, 100% of FusionAnalyser-versus-calliper differences fell within ±1.00mm (mean difference 0.20mm, 95% CI 0.12-0.29); agreement in the maxillary arch was marginally lower but clinically acceptable.
CONCLUSION
The free orthodontic software produced tooth-width and arch space deficiency measurements clinically comparable to those of the commercial platform and within clinically acceptable agreement limits of plaster-cast calliper measurements. Bolton ratio agreement was less interchangeable due to the cumulative aggregation of small per-tooth errors. These findings support the adoption of freely available orthodontics-specific software, particularly where commercial licensing is a barrier.
Unknown authors· International Orthodontics· 0 citations
Objective: To compare the diagnostic accuracy and time efficiency of a Standard Operating Protocol for Orthodontic Diagnosis (SOPOD) with routine diagnostic methods among dental interns.
Materials and Methods: Twenty seven dental interns with ≥2 weeks of orthodontic clinical posting participated. Baseline diagnoses were established by an experienced orthodontist. Each intern first diagnosed selected malocclusion cases using routine procedures based on Proffit’s approach, followed by re diagnosis of the same cases using SOPOD. Diagnostic accuracy scores and time taken were recorded.
Comparisons between methods were analyzed using paired t tests.
Results: SOPOD yielded significantly higher diagnostic accuracy and improved consistency compared with routine methods. Variability among interns was reduced and diagnostic time was shorter. Differences between the two approaches were statistically significant (p < 0.05).
Conclusion: SOPOD enhanced diagnostic accuracy and efficiency among dental interns. Incorporating SOPOD training into undergraduate orthodontic curricula may strengthen diagnostic competence and clinical decision making.
Niharika Khandelwal, J. Sodawala, Tanu Mahobia et al.· Journal of Dental Health and...· 0 citations
STATEMENT OF PROBLEM
Assessment of anterior dental and gingival esthetics has been commonly based on visual judgment and manual measurements. These approaches are time-consuming and show considerable examiner-dependent variation.
PURPOSE
The purpose of this study was to develop and validate a computer vision-based approach for automated extraction of anterior tooth morphologic parameters and the Pink Esthetic Score (PES) and the White Esthetic Score (WES) from intraoral scan data.
MATERIAL AND METHODS
Intraoral scans from 490 maxillary anterior teeth (245 pairs) of orthodontic patients were analyzed. An automated approach extracted dental and gingival contours, length-to-width ratios, and geometric and relative color differences (ΔE00, ΔL, Δa) compared with corresponding contralateral homologous teeth. Agreement between automated and manual measurements was evaluated using Bland-Altman analysis. The 245 pairs were randomly divided into exploration (n=175 pairs) and validation (n=70 pairs) sets. In the exploration set, logistic regression models combined with descriptive statistics established quantitative grading thresholds based on expert scores. In the validation set, agreement and reliability between automated and expert consensus scores were assessed using weighted Cohen kappa (κ) and intraclass correlation coefficients (ICC). Bland-Altman analysis and the Wilcoxon signed-rank test evaluated systematic bias. Furthermore, a Z test was performed to compare inter-examiner agreement with and without visual contour assistance across 245 pairs (α=.05).
RESULTS
Expert evaluations demonstrated moderate to almost perfect intra-examiner and inter-examiner agreement. The Bland-Altman analysis demonstrated excellent agreement between automated and manual measurements for tooth length-to-width ratios, indicating negligible systematic bias (mean difference=-0.001, 95% LoA: -0.061 to 0.058). Reliability between automated PES and WES scores and expert consensus scores was excellent for total PES (ICC=0.926, 95% CI: 0.881-0.954) and total WES (ICC=0.960, 95% CI: 0.936-0.975). Across all esthetic subcategories (95% CIs ranging from 0.639 to 1.000), agreement was almost perfect for 7 parameters (κ>0.890) and substantial for soft tissue contour (κ=0.790). No significant differences were found between automated and expert consensus scores (Wilcoxon P>.05). Furthermore, visual contour assistance improved inter-examiner agreement.
CONCLUSIONS
The developed approach demonstrated excellent agreement with expert consensus and provided stable, reproducible esthetic measurements from intraoral scan data. By integrating objective metrics with visual annotations, it offered a practical tool for standardized esthetic assessment and may support clinical decision-making in digitally assisted esthetic dentistry. Further validation in broader clinical settings is warranted.
Wenjia Chen, Tian Zhou, Mengyu Liang et al.· The Journal of prosthetic de...· 0 citations
A risk prediction model for post-clear aligner gingival embrasures was successfully developed and validated using multimodal oral data, with RF as the optimal algorithm that exhibits good discrimination, calibration, and clinical utility.
Haiyan Wang, Hanfei Shi, Liping Fan et al.· Acta Odontologica Scandinavi...· 0 citations
ABSTRACT Introduction: Mixed dentition analysis is essential for predicting space requirements and guiding early orthodontic decision-making. Moyers analysis is widely used due to its simplicity and clinical applicability, however, manual execution may increase execution time and introduce operator-dependent errors. Objective: This study aimed to develop and validate an automated digital tool based on Moyers analysis, to improve efficiency and support orthodontic education. Material and Methods: This cross-sectional study included 20 participants (10 undergraduate and 10 graduate dental students) who performed mixed dentition analysis on a standardized plaster model using the manual method and a software-based method at separate time points. Agreement between methods was evaluated using intraclass correlation coefficients (ICC), Bland-Altman analysis, and linear regression to assess proportional bias. Execution time was recorded and compared using paired t tests. Participant preference was also assessed. Results: Excellent reliability was observed (ICC > 0.95). No statistically significant differences were found between the manual and software-based methods (p > 0.05), and Bland-Altman analysis demonstrated substantial agreement without proportional bias. The software-based method reduced execution time by 38.2% (p < 0.001). All participants preferred the software-based method. Conclusion: The software-based method demonstrated agreement comparable to the manual Moyers analysis, while significantly reducing execution time. Its usability and efficiency support its application in orthodontic education and clinical workflow optimization.
Mariana Vasconcellos Bazoli Rodrigues, Luciana Rougemont Squeff, A. D. de Castro et al.· Dental Press Journal of Orth...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.