Aug 2026· ACM Transactions on Social Computing· 0 citations· 29 references
TL;DR
It is suggested that confidence interfaces should be evaluated for human–AI team behavior, with explicit support, transparency, and robustness checks, rather than for model-side statistical fidelity alone.
Abstract
As AI systems increasingly support consequential human decisions, the confidence they display shapes whether users accept or override their recommendations. A common assumption is that well-calibrated model confidence will induce well-calibrated human reliance. This paper shows that the two are distinct. It introduces calibrated reliance as a team-level property of human–AI decision making and proposes Reliance Calibration Error (RCE) as a metric for quantifying the gap between displayed confidence and realized team accuracy. Using 35,670 human–AI interactions from the HAIID benchmark and complementary evidence from the GRACE benchmark, the analysis identifies a systematic calibration–reliance gap: even near-calibrated models can produce miscalibrated reliance once confidence is interpreted through human judgment. The evidence is consistent with self-confidence moderating advice uptake, but the observational design does not identify anchoring as the unique mechanism. Because this gap arises at the interface between model confidence and human action, the paper develops a human-aware confidence communication framework that remaps displayed confidence without changing the underlying predictor. On held-out HAIID data, plug-in observed-outcome RCE indicates that subgroup-aware remapping can substantially reduce display–outcome misalignment, while model-predicted simulations show that bounded/global policies improve team MSE more reliably than aggressive subgroup remapping. Diagnostics also show that unconstrained remapping creates substantial boundary mass and that model-predicted counterfactual RCE is sensitive to aggressive display shifts. Bounded variants preserve much of the simulated decision-quality gain while avoiding 0/1 displays. A small prospective pilot is reported as a feasibility check rather than confirmatory validation. These findings suggest that confidence interfaces should be evaluated for human–AI team behavior, with explicit support, transparency, and robustness checks, rather than for model-side statistical fidelity alone. Code and materials for reproducing the analyses and using the confidence-display policies are available at https://github.com/OliverDOU776/From-Calibrated-Confidence-to-Calibrated-Reliance.
While participants rated hedged and unhedged AI as equally trustworthy and likely to be correct, they were significantly less likely to follow hedged advice in a binary choice, and how linguistic markers can be used to calibrate user reliance to model certainty is discussed.
Laura Spillner, Johanna Rockstroh, Nina Wenig et al.· International Conference on...· 0 citations
It is argued that cognitive misalignment represents a likely impediment to AI adoption in many envisioned applications, and that addressing it is important for creating AI systems on which users are both willing and justified to rely.
Vijay Keswani, Breanna K. Nguyen, Cyrus Cousins et al.· 0 citations
This work develops a six step BBN framework and illustrates it to model customer intention to consult a doctor in an alternative healthcare system and reveals that while self efficacy appears to be a major factor, its actual causal impact is small.
K. Rahul, Shovan Chowdhury Indian Institute of Management Kozhikode, Kerala et al.· 0 citations
Simulations show that over-reliance on a weak AI is especially harmful, and that diversifying AI signals across users can better keep the crowd informative, and conclude with implications for understanding human-AI interaction in information spread and designing misinformation interventions.
Zhuoran Lu, Weilong Wang, Yang-Yang Yu et al.· 0 citations
Claim-level decomposition combined with post-hoc calibration reduces expected calibration error on factual questions while exposing failure modes on adversarial false-premise questions where decision-makers most need reliable uncertainty estimates.
Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang et al.· 0 citations
Enterprise strategic decision support requires AI systems that are not only accurate, but also uncertainty-aware, risk-calibrated, explainable, and governance-compliant. This paper proposes TRUST-ESD, a risk-calibrated and governance-aware framework for enterprise decision support under uncertainty. TRUST-ESD evaluates feasible counterfactual strategies through predictive utility estimation, conformal uncertainty calibration, CVaR-based downside-risk scoring, risk-memory retrieval, policy-as-code governance, explainability, and human oversight. Unlike prediction-only methods that select actions by maximum expected utility, TRUST-ESD recommends strategies that balance value, reliability, risk exposure, and compliance. Experimental results show that TRUST-ESD improves risk-adjusted utility by 7.95%, reduces risk exposure by 23.22%, reduces CVaR by 23.78%, lowers calibration error by 13.89%, improves explanation fidelity by 10.90%, and increases governance compliance by 9.76% compared with strong uncertainty-aware baselines, while maintaining competitive predictive accuracy. Ablation and case-study analyses further confirm that uncertainty calibration, downside-risk scoring, risk memory, explainability, and governance validation jointly improve trustworthy enterprise decision-making.
Tian Qiu, Li Yan, Mahabubur Rahman Miraj et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.