Skip to content
Review Open access

Multimodal Large Language Models for Dental Chart Image Interpretation: Cross-Sectional Benchmarking Study With Students and Clinicians

Unknown authors
Sep 2026 · JMIR Medical Education · Vol 12 · 0 citations · 60 references
Medicine

Abstract

Abstract Background Entering clinical training, dental students must learn to read tooth-centered electronic dental records, but limited teaching time in patient-centered clinics can leave gaps in chart-reading literacy. Multimodal large language models (LLMs) that process dental record images may offer scalable educational support, yet their performance has not been benchmarked. Objective This study evaluated multimodal ChatGPT models on dental chart image interpretation and compared the best-performing model with dental students and residents. Methods A retrospective, cross-sectional benchmark study used deidentified dental chart images from 15 patients at Seoul National University Dental Hospital (2017‐2025). Charts containing Korean and English text were captured as sequential screenshots (154 images). For each patient, 24 Korean-language questions (360 total) spanned four categories: Type A, general factual retrieval; Type B, tooth- or procedure-specific retrieval; Type C, interpretation requiring multientry synthesis; and Type D, absent-information questions assessing abstention. Nine multimodal ChatGPT models (available August 2, 2025‐August 9, 2025) were evaluated under standardized conditions. Outputs were scored against a gold standard using 7 metrics, with sentence bidirectional encoder representations from transformers (SBERT) similarity prespecified as the primary semantic measure. Human baselines included 2 third-year students and 2 first-year residents. Groups were compared with Kruskal-Wallis tests and Dunn post hoc analyses. Gold-standard reliability was assessed by independent senior-expert review and chance-corrected agreement (Gwet AC1) among clinical reference raters, and the model ranking was confirmed by content-based clinical accuracy analysis. Results Across 360 items, GPT-5 Thinking achieved the highest median SBERT similarity (0.900, IQR 0.525‐1.000), followed by GPT-5 Pro (0.861, IQR 0.501‐1.000) and OpenAI o3 (0.831, IQR 0.489‐1.000), with significant overall group differences (P<.001). Student 1 did not differ significantly from GPT-5 Thinking across Types A-D (all P≥.16), whereas Student 2 differed only on Type D (P=.002; Cliff δ=−0.27). Resident 1 scored higher than GPT-5 Thinking on Types A (P=.01) and C (P=.04) but lower on Type D (P<.001; δ=−0.35), whereas Resident 2 scored higher on Types A (P=.03), B (P=.02), and C (P=.04) and did not differ on Type D (P=.23). Type C tasks showed compressed SBERT distributions and low exact-match rates, indicating persistent difficulty in synthesis. The clinical reference rater agreement was high (Gwet AC1=0.99), and the content-based clinical-accuracy ranking was closely aligned with the SBERT ranking (Spearman ρ=0.90; P=.001). Conclusions Statistically nonsignificant differences were observed between GPT-5 Thinking and dental students for most question-type contrasts, with Student 2 differing only on absent-information items. Compared with first-year residents, GPT-5 Thinking remained lower on several Type A-C contrasts, particularly interpretive Type C and one tooth- or procedure-specific Type B comparison; Type D contrasts require cautious interpretation because all groups had ceiling medians. Within this single-center benchmark, multimodal LLMs may have potential as supervised educational tools for chart-reading practice and verification, rather than as replacements for clinical expertise, pending external validation across institutions, specialties, and record systems.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.