Skip to content
#federated learning Open access

Beyond Diagnostic Accuracy: Calibration and Predictive Uncertainty of Artificial Intelligence Models for Oral and Dental Disease Diagnosis

Oct 2026 · Figshare
Artificial Intelligence in Healthcare and Education

Abstract

A Systematic Review Protocol “Beyond Diagnostic Accuracy: Calibration and Predictive Uncertainty of Artificial Intelligence Models for Oral and Dental Disease Diagnosis” Title of the ReviewBeyond Diagnostic Accuracy: Calibration and Predictive Uncertainty of Artificial Intelligence Models for Oral and Dental Disease Diagnosis Background and ReasoningArtificial intelligence (AI), machine learning (ML), and deep learning (DL) are increasingly used for diagnostic applications in oral medicine and dentistry, including oral cancer screening, oral potentially malignant disorder recognition, epithelial dysplasia assessment, dental caries detection, jaw-lesion diagnosis, radiographic disease classification, and multimodal oral lesion assessment. Most diagnostic AI studies are evaluated primarily with discrimination metrics such as accuracy, sensitivity, specificity, F1-score, and area under the receiver operating characteristic curve (AUROC). Although these measures indicate how well a model separates diagnostic categories, they do not establish whether the probabilities assigned to predictions are reliable.Probability calibration evaluates agreement between predicted probabilities and observed outcome frequencies. Relevant measures include the Brier score, expected calibration error (ECE), maximum calibration error (MCE), calibration or reliability plots, calibration slope and intercept, calibration-in-the-large, and observed-to-predicted ratios. Post-hoc approaches such as Platt scaling, isotonic regression, and temperature scaling may be used to improve probability calibration.Predictive uncertainty is related to, but conceptually distinct from, calibration. Uncertainty quantification (UQ) attempts to identify individual predictions for which the model may be unreliable. Approaches include Bayesian neural networks, Monte Carlo dropout, deep ensembles, predictive entropy or variance, and explicit decomposition of epistemic and aleatoric uncertainty. Clinically, these estimates may support selective prediction, abstention, specialist referral, additional imaging, or biopsy rather than forcing an automated prediction for every case.A systematic review focused specifically on calibration and prediction-level uncertainty is therefore warranted to determine how these reliability dimensions are evaluated in oral and dental diagnostic AI, whether they are assessed under external validation or distribution shift, and whether they are translated into clinically actionable decision-support strategies.Objectives• Identify diagnostic AI studies in oral and dental disease that quantitatively assess probability calibration and/or prediction level uncertainty.• Identify methods used to quantify predictive uncertainty, including epistemic and aleatoric uncertainty.• Determine whether calibration and uncertainty are evaluated under independent external validation, population shift, technical shift, or other forms of distribution shift.Review Questions· How are probability calibration and predictive uncertainty assessed in AI models for diagnosing oral and dental diseases, and what evidence shows that these methods improve the trustworthiness and safety of diagnostic forecasts?· What measures and techniques are employed to assess probability calibration?· What techniques are employed to measure prediction-level uncertainty?Eligibility CriteriaInclusion Criteria· Primary empirical studies involving human oral or dental disease, human-derived clinical cases, oral/dental images, radiographs, lesions, teeth, or other clinically relevant oral/dental diagnostic data.· Studies evaluating AI, ML, DL, Bayesian models, ensemble models, federated-learning models, multimodal/foundation models, large language or vision-language models, or computer-aided diagnostic systems.· Diagnostic, detection, classification, grading, differential-diagnosis, or disease-screening tasks.Studies that quantitatively evaluate probability calibration and/or prediction-level uncertainty.· Eligible calibration evidence includes the Brier score, ECE, MCE, calibration/reliability plots, calibration slope/intercept, calibration-in-the-large, observed-to-predicted measures, or an explicit probability-recalibration procedure.· Eligible predictive-UQ evidence includes Bayesian inference, MC dropout, predictive variance or entropy, ensemble disagreement, epistemic uncertainty, aleatoric uncertainty, or another quantitative prediction-level uncertainty measure.· English language studies published from 1 January 2010 to 4th September 2026.Exclusion Criteria· Studies unrelated to oral or dental disease.· Non-English publications.· Only studies focused on prognosis or future risk prediction are included, provided they do not involve a current diagnostic, detection, classification, grading, or screening task.· Pure anatomical or lesion segmentation studies unless uncertainty is evaluated at the disease-diagnostic level or explicitly used for disease-level clinical triage.· Active learning, uncertainty-aware representation learning, or boundary-learning studies that do not quantify prediction-level diagnostic uncertainty.· Studies reporting only confidence intervals around accuracy, AUROC, sensitivity, specificity, or other conventional performance metrics; such intervals are not considered predictive uncertainty.· Animal, preclinical-only, or purely synthetic studies without clinically relevant human evaluation.· Reviews, editorials, commentaries, protocols, letters without primary empirical results, and abstracts lacking sufficient methodological or outcome information. Data Sources and Search StrategyDatabases: Pubmed, Scopus, and IEEE XploreSearch Terms: "Oral" OR "Dental" OR "Dentistry" OR "Caries" OR "Periodontal" OR "oral cancer" OR "Oral lesion" OR "Oral potentially malignant disorder" AND"Artificial Intelligence" OR "Machine Learning" OR "Deep Learning" OR "Neural Network" OR "Large Language Model" OR "Computer Aided Diagnosis"AND"calibration" OR "Brier" OR "expected calibration error" OR "uncertainty" OR "Bayesian" OR "Monte Carlo dropout" OR "epistemic" OR "aleatoric" OR "selective prediction" OR "abstention" Study Selection ProcessAll retrieved records will be consolidated and duplicates removed before screening. Titles and abstracts will be assessed against the predefined eligibility criteria, followed by full text assessment of potentially eligible studies. Reasons for exclusion at the full-text stage will be recorded. The selection process will be summarized using a PRISMA 2020 flow diagram.Reviewer procedure: For a publication-quality protocol, title/abstract screening and full text eligibility assessment should be performed independently by two reviewers, with disagreements resolved by discussion or a third reviewer. This statement should be retained only if that procedure is actually implemented; otherwise the protocol and manuscript must describe the true screening process. Data ExtractionA standardized extraction form will be used. Missing or unreported information will be coded as Not Reported (NR) and will not be inferred. The extraction framework will include the following variables:· Bibliographic information: author, year, country, publication type, journal/conference.· Clinical characteristics: disease/domain, diagnostic target, setting, study design, sample size, unit of analysis, class distribution, modality, and reference standard.· AI methodology: architecture/model, feature strategy, training strategy, data split/cross-validation, independent test set, external cohort, multicenter status, human comparator, explainability, and code/data availability.· Discrimination: accuracy, balanced accuracy, sensitivity, specificity, precision, F1-score, AUROC, AUPRC, and other reported diagnostic metrics.· Calibration: Brier score, ECE, MCE, reliability/calibration plots, observed-to-predicted measures, calibration slope/intercept, calibration-in-the-large, and recalibration methods.· Predictive uncertainty: Bayesian methods, MC dropout, ensembles, predictive entropy/variance, epistemic uncertainty, aleatoric uncertainty, or other prediction-level UQ.· Study limitations and methodological concerns.OutcomesPrimary OutcomesThe primary outcomes are measures of diagnostic probability reliability and prediction-level uncertainty. Calibration outcomes include Brier score, ECE, MCE, calibration/reliability plots, calibration slope and intercept, calibration-in-the-large, observed-to-expected or observed-to-predicted measures, and changes after recalibration. Predictive-UQ outcomes include predictive entropy or variance, Bayesian posterior/predictive uncertainty, MC-dropout uncertainty, ensemble disagreement/variance, and explicit epistemic and/or aleatoric uncertainty.Secondary outcomesSecondary outcomes include conventional diagnostic discrimination metrics and measures of reliability translation, including external calibration, domain-shift performance, OOD detection, selective prediction, abstention/referral rate, risk-coverage/selective-risk analysis, uncertainty-error association, human-AI comparison, decision-curve analysis, net benefit, clinical-loss analysis, and other measures of uncertainty-guided clinical utility.Risk of Bias and Methodological QualityMethodological quality will be evaluated using a PROBAST+AI-informed framework appropriate to AI-based prediction and diagnostic models. The principal domains will include participants/data sources, predictors/input data, outcome/reference standard, and analysis/model evaluation.AI-specific considerations will include independence of development and evaluation data, patient-level versus image-level separation, potential data leakage, effective sample size relative to model complexity, hyperparameter optimization, handling of class imbalance, probability-c

View source

Similar papers

#machine learning Review Open access Oct 2014

Software development in startup companies: A systematic mapping study

The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.

Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al. · 394 citations · ⚡54
#machine learning Review Open access Jun 2014

Why Early-Stage Software Startups Fail: A Behavioral Framework

This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.

Carmine Giardino, Xiaofeng Wang, P. Abrahamsson · 175 citations · ⚡19
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#machine learning Review Open access May 2016

Key Challenges in Software Startups Across Life Cycle Stages

It is found that what perceived as biggest challenges by software startups do vary across different life cycle stages, even though its significance decreases when the learning focuses of the startups move from problem to solution and their products mature.

Xiaofeng Wang, Henry Edison, Sohaib Shahid Bajwa et al. · 62 citations · ⚡6

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.