Skip to content
Open access

A multimodal vision-language model for comprehensive dental diagnosis and enhanced clinical practice

Jul 2026 · Nature Communications · Vol 17 · 1 citation · 79 references
Medicine

TL;DR

DentVLM is a dental vision-language model developed to support dental diagnosis across seven oral imaging modalities and 36 tasks that surpasses junior readers, matches intermediate general practitioners and approaches senior specialists, and reduces diagnostic time by 15.0-37.0% in collaborative clinical workflows.

Abstract

Oral diseases affect billions of people, yet specialist dental expertise remains unevenly distributed, and diagnosis often requires synthesis across diverse imaging modalities. Existing artificial intelligence systems mostly address isolated tasks, limiting their applicability in comprehensive dental assessment. Here we introduce DentVLM, a dental vision-language model that jointly interprets images and text, supports expert-level oral disease diagnosis across seven dental imaging modalities and 36 tasks. Developed using 110,447 images and 2.46 million bilingual visual question-answer pairs, DentVLM outperforms leading proprietary, open-source and domain-specific medical models on internal and external tests. In a study of 32 participants, DentVLM surpasses junior readers, matches intermediate general practitioners and approaches senior specialists. In collaborative workflows, it raises junior and intermediate readers toward specialist-level performance and reduces diagnostic time for all readers by 15.0-37.0%. These results establish DentVLM as a clinical decision support tool for reducing specialist care gaps and broadening access to high-quality dental expertise. DentVLM is a dental vision-language model developed to support dental diagnosis across seven oral imaging modalities and 36 tasks. It matches intermediate general practitioners, approaches senior specialists, and reduces diagnostic time by 15.0-37.0% in collaborative clinical workflows.

Read PDF

Similar papers

Review Open access Jul 2026

Artificial intelligence for dental caries diagnosis: translating algorithms to clinical practice

Dental caries remains the most prevalent chronic disease worldwide, affecting more than two billion people and driving substantial healthcare costs. While early detection is central to minimally invasive dentistry, traditional diagnostic methods suffer from limited sensitivity and high inter-examiner variability, especially for incipient lesions. Artificial Intelligence (AI) has shown promise in overcoming these limitations, and many studies have reported expert-level performance under controlled conditions. However, a substantial translational gap persists between algorithmic success in silico and reliable performance in real-world clinical environments. This review synthesizes the full development pipeline of AI for caries diagnosis—from data curation and ground-truth construction to model design, validation, and clinical deployment. We highlight persistent bottlenecks including domain shift across imaging devices and clinical settings, subjective and inconsistent annotation practices, limited multimodal datasets, and heterogeneous reporting standards. Emerging strategies such as multi-center data collection, probabilistic labeling, self-supervised learning, domain adaptation, and test-time augmentation offer partial solutions but remain underutilized. We argue for a paradigm shift from binary detection toward quantitative, risk-based staging that aligns with minimally invasive dentistry and the WHO Global Oral Health Action Plan 2023–2030. By advocating for standardized multimodal datasets, rigorous external validation, explainable interfaces, and human-centered clinical integration, this review outlines a roadmap for translating AI innovation into trustworthy, equitable, and clinically meaningful decision-support systems capable of reducing the global burden of untreated caries.

Ziyu Wang, Zixuan Zhu, Ti Jiang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning

DentAgent is introduced, an evidence-centric multi-agent framework, in which the Orchestrator coordinate five specialized agents spanning various modalities, which supports its value for broadly applicable and traceable multimodal dental reasoning, and highlights its potential as a technical foundation for population oral health assessment and management.

Zijie Meng, Xi-Wei Dai, Yixuan Tang et al. · 0 citations
Open access Jul 2026

Detecting Dental Caries Using General-Purpose Large Multimodal Models From Oral Photographs.

OBJECTIVES To determine whether a general-purpose large multimodal model (LMM) can detect dental caries from intraoral photographs without task-specific training, evaluating image-level classification and tooth-level localisation under zero-shot (no reference examples) and five-shot (five annotated examples) conditions. METHODS This diagnostic accuracy study used the benchmark test split of 1255 publicly available intraoral photographs. Gemini 3.1 reasoning and instant were queried via Vertex AI API using stateless calls and structured prompts for sequential, tooth-by-tooth scanning. Each model under different configurations was evaluated across 10 independent inference runs. Predictions (bounding boxes) were evaluated against expert annotations using greedy Intersection-over-Union (IoU ≥ 0.5). True positives required spatial overlap at the tooth level, or at least one correctly localised lesion at the image level. Performance was summarised using sensitivity, precision, F1-score, and mAP@50. RESULTS Performance was strongly dependent on image view and model type, with reliable results observed mainly for the reasoning models on occlusal images. At the image level, zero-shot prompting showed high sensitivity but lower precision, yielding values of 0.95, 0.80, and 0.87 for sensitivity, precision, and F1-score, respectively. Five-shot prompting produced a more balanced profile, with corresponding values of 0.88, 0.87, and 0.87. At the tooth level, zero-shot reasoning showed a similar sensitivity-prioritised pattern, with sensitivity, precision, F1-score, and mAP@50 of 0.88, 0.49, 0.63, and 0.62, respectively. Five-shot prompting improved precision and produced a more balanced localisation profile, with corresponding values of 0.77, 0.57, 0.64, and 0.63. CONCLUSIONS General-purpose LMMs such as Gemini show potential as a scalable, automated caries screening tool for teledentistry without domain-specific fine-tuning. Although prompt engineering effectively modulates the sensitivity-precision trade-off, the model's tendency toward over-detection requires rigorous clinical validation and governance before real-world deployment as a public screening tool.

M. Moharrami, Sina Asadi, Owais A. Farooqi et al. · 0 citations
Preprint Aug 2026

Open-Linguistic Concept Unified Learning for Cross-Site Interpretable Dermatology Image Diagnosis

Human-interpretable computer-aided diagnosis is crucial for clinical decision making. Concept-based models excel by providing transparent reasoning and enabling post-hoc, clinician-in-the-loop interventions. However, their rigid dataset-specific adaptation inherently restricts cross-site generalization. Applying them across diverse modalities, such as dermoscopic and clinical photographs, is challenging due to heterogeneous concept taxonomies varying in availability, granularity, and semantics across cohorts. Consequently, adapting Foundation Vision-Language Models (FVLMs) demands costly label engineering and repeated post-training. Existing intervention mechanisms remain rigidly tied to predefined concepts, lacking adaptability and hindering scalable dermatology CAD deployment. To address these bottlenecks, we propose UniCon, an open-linguistic unified concept learning framework for multimodal interpretable vision-language diagnosis. UniCon resolves these challenges through three contributions: (1) A shared semantic representation space via a unified concept prototype codebook, seamlessly coordinating heterogeneous concept systems across modalities without dataset-specific retraining. (2) Open-linguistic based multi-faceted semantic specifications to overcome sparse textual label limitations, improving boundary sensitivity in uncertain clinical contexts. (3) A robust, cross-site adjustable intervention interface powered by reliability-gated bottleneck aggregation, enabling consistent reasoning and transferable clinician corrections. Extensive experiments demonstrate that beyond securing top-tier diagnostic accuracy, UniCon successfully bridges disparate clinical taxonomies, unlocking unprecedented cross-site intervention capabilities. Code is available at https://github.com/wuchengyu123/UniCon.

Chengyu Wu, Junpeng Tan, Wanxiang Luo et al. · 0 citations
Review Open access Aug 2026

Vision and Language Models for Classifying Maxillary Sinus Disease on Cone-Beam Computed Tomography: A Transparent Multimodal Benchmark

Background: Cone-beam computed tomography (CBCT) frequently captures the maxillary sinuses incidentally, and reliable automated detection of sinus abnormality is clinically relevant. Unlike most vision-language benchmarks in medical imaging, which pair images with pre-existing, human-authored clinical reports, findings text can also be generated directly by a large language model from the image itself--raising the question of how much diagnostic value such AI-derived text carries, and whether that value depends on independent verification. Multimodal artificial intelligence (AI) benchmarks risk overstating performance if the provenance of each input--image, raw AI-generated text, or radiologist-verified text--is not clearly separated and reported. Methods: We used 300 mid-sagittal CBCT slices from the MMDental dataset. ChatGPT generated findings text and a provisional normal/abnormal label for every slice (majority vote, three independent readings from the image alone); primary classification performance was assessed on this full, unfiltered set (n=300). A radiologist then independently reviewed each case's image together with ChatGPT's description, producing their own diagnosis; three cases were excluded as insufficient, yielding 297 confirmed cases. On this subset, every model was retrained and re-evaluated under identical 10-fold cross-validation on both the provisional ChatGPT-only labels ("pre") and the radiologist-confirmed labels ("post"), isolating the effect of label provenance from image or architecture. Eight vision architectures, seven language classifiers, and five VLMs were evaluated throughout; three generative models performed exploratory note-drafting. Findings: Raw ChatGPT-generated text produced the highest performance of any modality or condition: language models reached near-ceiling AUC (0.992 to 1.000, n=300), exceeding every vision model (AUC 0.799 to 0.880) and every VLM image-only probe (AUC 0.63 to 0.69). On the 297-case pre/post analysis, this advantage depended heavily on label source: language and text-derived VLM performance fell substantially from ChatGPT-only to radiologist-confirmed labels (e.g. BERT-base AUC 0.999 to 0.837), while vision-model performance was stable or modestly improved (e.g. DenseNet-121 0.867 to 0.891). The radiologist reclassified 62 of 297 cases (21%) relative to ChatGPT's provisional read, and a meaningful proportion of raw ChatGPT text was clinically uninterpretable or unsupported by the imaging. Interpretation: As shown here for the first time, raw, image-derived AI-generated text yields the highest apparent classification performance in this benchmark, but this reflects the text's alignment with its own self-generated labels rather than verified diagnostic content, and a substantial share of that text is not clinically explainable. Radiologist-confirmed text and labels give a lower but trustworthy estimate of true performance, on which convolutional neural network (CNN) vision models remain a stable, comparatively inexpensive baseline. Multimodal dental AI should report performance separately by modality and label provenance rather than pooling headline metrics.

S. Alhebshi, H. Khalifa, T. D. Pham · 0 citations
#large language models Review Open access Sep 2026

Multimodal medical diagnosis: a mini review of LLM–vision fusion models in low-resource healthcare settings

Recent advances in large language models (LLMs) and vision transformers have enabled multimodal systems that integrate clinical text with medical imaging for diagnostic decision-making. While these systems show promising results on benchmark datasets in well-resourced research settings, their applicability in low-resource healthcare environments where diagnostic disparities are most severe remains limited and poorly understood. This mini review synthesizes key developments in LLM–vision fusion architectures from 2018 to 2026, with a focus on radiology-oriented visual question answering (VQA) and report generation systems viewed from a deployment perspective. Rather than comprehensively cataloguing multimodal medical AI, we synthesize the evolution of LLM–vision fusion architectures and discuss complementary deployment-enabling strategies, including parameter-efficient adaptation, post-training quantization, federated learning, and multilingual support, where they directly improve the feasibility of radiology AI in resource-constrained healthcare settings. Rather than focusing solely on performance benchmarks, we examine these approaches through a deployment-oriented lens, highlighting trade-offs between representational capacity, computational efficiency, interpretability, and memory footprint. We argue that current progress remains substantially shaped by model scaling and benchmark optimization, which often do not address the memory, connectivity, and annotation constraints of low-resource healthcare systems. While cross-modal transformer architectures provide strong representational alignment, their computational demands and reliance on large curated datasets limit real-world deployment. In contrast, emerging directions including parameter-efficient fine-tuning, post-training quantization, federated learning, and modular agent-based systems offer more tractable pathways toward clinical integration under hardware and data constraints. To bridge the gap between benchmark performance and clinical utility, we identify concrete challenges in data scarcity, multilingual coverage, and calibration, and propose a shift toward lightweight, interpretable, and hardware-aware multimodal AI. This perspective highlights the need to move beyond scaling-centric design toward models that can run on 4–8 GB VRAM, operate offline, and generalize across languages and imaging equipment.

Kahakashan Ashraf, Md.Hamid Hosen, N. Farah et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.