Validation of a commercial intraoral auto-segmentation algorithm against expert-annotated clinical data with comparison to a transparent algorithm
Abstract
Artificial intelligence (AI)-based tooth segmentation has the potential to improve the efficiency and consistency of digital clinical workflows; however, the generalizability of commercially available algorithms on clinical datasets and their interpretability remain limited. This study proposed a two-level evaluation method. It evaluated the performance of a black-box commercial algorithm using an expert-annotated clinical reference dataset of 126 intraoral scans representing diverse clinical presentations, including severe crowding, spacing, missing teeth, and dental restorations. Two specialists independently generated reference annotations. In the next level, a transparent two-stage segmentation method combining YOLOv8-based tooth localization and numbering with a 3D U-Net segmentation model was developed using 80 scans from an open-source dataset and evaluated alongside the commercial algorithm. Performance was evaluated using the Dice Similarity Coefficient (DSC) and tooth classification accuracy. Tooth detection and numbering were assessed descriptively using confusion matrices and accuracy. End-to-end segmentation performance was compared at the scan level using paired parametric or nonparametric tests, depending on the distribution of paired differences. Regional performance differences between anterior and posterior regions and between maxillary and mandibular jaws were assessed using independent parametric or nonparametric tests, as appropriate. The commercial algorithm achieved significantly higher segmentation performance than the transparent model, with a median DSC of 0.90 (IQR 0.86–0.95) versus 0.76 (IQR 0.66–0.84), respectively ( p < 0.001). In contrast, the transparent model showed tooth-numbering accuracy of 0.92 versus 0.83. No significant differences in DSC were observed between anterior and posterior regions for either algorithm. Maxillary and mandibular performance did not differ significantly for the commercial algorithm, whereas the transparent model performed significantly better in the mandible. The commercial algorithm demonstrated strong segmentation performance on an independent, expert-annotated clinical dataset. Although the transparent approach achieved lower segmentation accuracy, it offered greater methodological transparency and insight into decision-making.