The Robust Ambiguity Detection (RAD) framework is advanced for quantifying predictive ambiguity using two complementary metrics: Model-Space Consistency and Feature-Space Consistency, which provide an interpretable characterisation of the sources of ambiguity and the actions a user may consider in response.
Abstract
Machine learning models should be robust, in the sense of remaining predictively consistent under permissible variations. A model's predictions should ideally remain unchanged when it is replaced by a functionally equivalent one, or when its inputs are subject to minor, admissible perturbations. If such changes alter a prediction significantly, then the prediction is"ambiguous"with respect to the model. Models should abstain from making such ambiguous predictions and/or should flag them for human inspection, especially in high-stakes decision-making scenarios. However, in practice, such ambiguity is not easy to identify once a model is deployed. Here, the Robust Ambiguity Detection (RAD) framework is advanced for quantifying predictive ambiguity using two complementary metrics: Model-Space Consistency and Feature-Space Consistency. These two scores, the RAD Score-Pair, visualised through the RAD Plot, provide an interpretable characterisation of the sources of ambiguity and the actions a user may consider in response. RAD is evaluated on synthetic datasets with systematically controlled overlap, as well as several real-world datasets where the level of ambiguity cannot be directly inspected. Finally, we demonstrate a downstream application of RAD where samples are ranked by their RAD Pareto-Rank and the most ambiguous are abstained from prediction, achieving performance comparable to existing rejection-based approaches.
This work studies multiplicity from the perspective of auditing incorrect ensemble predictions, where the decision to divert an instance for human review is based on a consistency criterion that combines the ensemble margin with a measure of local prediction variability for each constituent model.
Sinjini Banerjee, Tim Marrinan, Anand D. Sarwate· 0 citations
Concept-Residual eXpansion (CRX), a concept-augmented framework that improves robustness by expanding the set of candidate predictive features by improving robustness to spurious correlations, is proposed.
Eric Xie, Guang-Zhi Xiong, Wenqian Ye et al.· Proceedings of the 32nd ACM...· 0 citations
Maximum likelihood (ML) estimation is a principled and statistically efficient approach for learning probabilistic models. However, for unnormalized models, ML estimation requires evaluating the partition function and differentiating through it, which may not always be tractable. Score matching provides a practically v...
Nishanth Shetty, Saisuchith Mahajan, C. Seelamantula· 0 citations
The Bayes factor (BF) is a central tool in Bayesian hypothesis testing and model selection, yet its practical use is often challenged. Classical BFs depend heavily on prior specification, cannot be applied with improper priors, and are typically interpreted through arbitrary evidence scales. Moreover, they fail to capt...
A probabilistic binary classifier is judged almost everywhere by discrimination - accuracy, the ROC curve, the area under it. Every such criterion is invariant to a monotone distortion of the predicted probabilities, so a classifier can rank perfectly and still return probabilities that are badly wrong. Calibration is...
Ebrahim Khaled Ebrahim, Ahmed El-Kotory· 1 citation
Machine-learning models are commonly developed under an assumption that training and test data are sufficiently complete, balanced, labelled, and drawn from compatible distributions. In practice, one or more of these conditions is often violated. Measurements may be missing or corrupted, rare classes may be poorly repr...
Masoumeh Zareapoor· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.