This work shows that a distance-minimization-based CE is mathematically equivalent to the maximum a posteriori estimate of a Gibbs posterior within the generalized Bayes framework, and introduces two decision rules beyond MAP within a unified framework.
Abstract
Counterfactual explanations (CEs) enhance the interpretability of machine learning models by identifying the smallest change to an input required to obtain a desired output. Although CEs are conventionally formulated as a distance-minimization problem, the theoretical basis of this formulation has received limited attention. We show that a distance-minimization-based CE is mathematically equivalent to the maximum a posteriori (MAP) estimate of a Gibbs posterior within the generalized Bayes framework, specifically when a distance-based prior is used. We call this formulation the Distance-Prior Generalized Bayes CE (DP-GBCE). Building on this posterior perspective, we introduce two decision rules beyond MAP within a unified framework: a Bayes decision that minimizes expected decision loss and CVaR-CE, a risk-averse decision rule. We also propose an extension that uses Bayesian model weights to mix the posterior distributions of multiple models, thereby accounting for model multiplicity, where several models have comparable predictive performance. Finally, we define metrics for evaluating both individual CEs and the posterior distribution as a whole, and use experiments on simulated data and Google Trends data to quantify the trade-offs among the decision rules.
The Bayes factor (BF) is a central tool in Bayesian hypothesis testing and model selection, yet its practical use is often challenged. Classical BFs depend heavily on prior specification, cannot be applied with improper priors, and are typically interpreted through arbitrary evidence scales. Moreover, they fail to capture uncertainty inherent in the data, leading to an analogy with frequentist p-values, and primarily reflect prior-predictive rather than posterior-predictive performance. We introduce the stochastic Bayes factor (SBF), a new framework that extends the BF by explicitly incorporating uncertainty via replicated data. Formally, the SBF is defined as a push-forward measure transferring the BF from the observed data space to that of replications. This approach generalizes previous calibration proposals, while emphasizing posterior-predictive replication as a robust alternative. We establish key theoretical properties, including model consistency, compatibility and dominance, ensuring that SBFs preserve desirable Bayesian guarantees. An algorithmic routine is then proposed to operationalize the SBF, guiding model discrimination in a principled way while naturally providing model calibration. Simulation studies and real applications confirm that the SBF offers improved robustness and predictive reliability compared to the classical BF, by providing a valuable tool for model comparison.
Bayesian inference for A/B testing is a family of prior and stopping-rule configurations with fundamentally different statistical properties, but it is often discussed as a single method, and no systematic overview exists. This paper organizes common configurations into a three-tier hierarchy: 1) posterior coherence with no error control, 2) false positive rates bounded under continuous monitoring via Bayes factor stopping, and 3) false discovery rate control and calibrated shrinkage via empirical Bayes. Many commercial platforms operate at the lowest tier by default. We show that Bayes factor stopping is near-optimal for a broad class of cost functions, including most proposed in the A/B testing literature; because the same rule also controls the false positive rate, the choice between a decision-theoretic and a frequentist formulation is largely one of parameterization. Furthermore, the empirical Bayes prior is the only path to the third tier, but winner-selected corpora, pooled programs, and heterogeneous metrics can each prevent calibration regardless of corpus size. Simulations against group-sequential and always-valid frequentist baselines show that flat-prior posterior stopping exactly reproduces naive peeking, that a well-calibrated empirical Bayes prior achieves the lowest estimation error, and that expected-loss stopping minimizes regret only when shipping a null-effect variant is nearly free. Error rates, estimation accuracy, and regret are all different risks, and the appropriate method follows from the risks an experimentation program needs to control, not the other way around.
A Bayes-filtered transformer (BFT) is a transformer trained on sequences that are generated in two steps: first a latent task is drawn from a prior, then observations are drawn conditional on that task. Trained under autoregressive log loss, the BFT's next-token prediction, in the idealized limit, is the Bayesian posterior predictive distribution (PPD) induced by that prior and that conditional law. In practice the trained BFT is only an approximation of this ideal PPD, raising an interpretive question: what prior and posterior over the latent task has the trained BFT actually internalized? Existing work answers this question by comparing the trained BFT's predictions against the predictions of various"reference"posteriors, each standing in for a different candidate algorithm or computation the BFT might be implementing. This prediction-space comparison is fragile: different posteriors can share the same posterior-mean predictions. We use predictive Monte Carlo (PMC) as a general interpretability tool for any BFT: using only next-token generation, PMC returns an approximation to the implicit prior and posterior over the latent task, answering the interpretive question directly in latent space. We apply PMC to three stylized task families spanning 0-Markov and 1-Markov exchangeability. The phenomena previously reported in these settings remain visible in latent space. Code is available at https://github.com/afiq-aswadi/bft-pmc
In modern Bayesian computation and parametric estimation, the marginal likelihood, serving as the denominator P(D) in Bayes'Theorem, is routinely bypassed via unnormalized proportionality relations. Even within specialized model-selection frameworks where it is explicitly evaluated to compute Bayes Factors, the denominator is treated purely as a static constant. This note evaluates a subtle analytical oversight resulting from this computational convenience. Through the analysis of a simplified, sequential partial-information system, we show that the marginal probability possesses a critical dual layer of information: while the posterior probability determines the local magnitude of a belief update upon a solitary trial, the marginal denominator governs the physical, long-run frequentist cadence of that update across a historical horizon. Discarding the denominator computes what an observer ought to believe once a specific dataset manifests, but erases the data-generating reality that governs how frequently that inferential state occurs in nature. We propose reclaiming the marginal likelihood as an active, real-time regularizer. We introduce three diagnostic measures to regulate recursive online estimation gains, and construct valid mixture probabilities that blend the prior and posterior to function as surprise-activated or conservative regulators. This paired probability framework may offer a robust regularizing mechanism for sequential estimation architectures and quantitative risk management scenarios under non-stationary distribution shifts.
This paper proposes the Bayesian Expected Uncertainty Reduction (B-EUR) model, which formalizes the value of trying a candidate design action as its expected reduction of epistemic uncertainty about action--outcome relations. The model addresses one part of the Uncertainty Driven Action (UDA) model's open question concerning how changes in uncertainty perception determine action selection. We examine two environmental properties: generalizability, or how far knowledge from one trial extends to neighboring candidates, and outcome discriminability, or how clearly differences among outcomes can be distinguished. We tested the model through simulations and human experiments using a graph-shape guessing task that isolates learning about action--outcome relations under a limited trial budget. Epistemic value followed an inverted-U-shaped relationship with generalizability and increased with outcome discriminability in the simulations. In the human experiments, the subjective value of trying and enjoyment followed inverted-U-shaped relationships with generalizability, while choice behavior reflected both properties. The B-EUR model provides a computational account of candidate-action evaluation within uncertainty-driven design activity and offers implications for constructing prototype sets, framing design problems, and organizing feedback to support informative exploration.
Shimon Honda, Takuma Miyaguchi, Koji Koizumi et al.· 0 citations
We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axioms and argumentation, but instead consequences, of higher-level modelling decisions. We show how epistemic and aleatoric uncertainty measures can be derived via decomposition of a subjective risk, based on a strictly proper loss. Reverse cross entropy provides a prominent example, where decomposition recovers the classic information-theoretic uncertainty terms. The same approach recovers numerous measures previously proposed across the UQ literature, providing them a common theoretical foundation. This suggests a new approach to UQ: given a modelling scenario and strictly proper loss, the corresponding epistemic and aleatoric terms are induced by the subjective-risk decomposition. We then extend our view to learning theory: we introduce and analyse subjective risk analogues of excess risk, approximation error and estimation error, and identify the connections to UQ. We consider this a first step towards a full learning-theoretic framework for uncertainty quantification.
R. Alamri, Michele Caprio, Gavin Brown· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.