RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty, and proves a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates.
Abstract
How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), which ties caution to epistemic uncertainty: the agent holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein ambiguity set whose radius is a monotone function of that posterior. The radius contracts with evidence, so behaviour interpolates continuously between worst-case robustness and risk-neutral total-reward maximization. The design follows the duality underlying the Entropic Value-at-Risk, which converts the choice of a risk level into the choice of an ambiguity radius. We show the resulting planning problem is well posed under transience and compactness conditions, and prove a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates. In a canonical binary-hazard instance, the induced criterion reduces to Conditional Value-at-Risk at a level set by the posterior entropy. A worked example shows the agent deferring the efficient action until a sharp identification threshold. RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty.
An agent still learning its environment should be cautious while ignorant and bold once confident. The entropic value-at-risk captures this through a robust-optimization identity---a confidence level fixes the radius of a relative-entropy ball of alternative models---but that ball cannot reach catastrophes the nominal deems impossible, precisely what a safe agent must hedge. We instead use an optimal-transport ball and study the coherent risk measure it induces, the Wasserstein entropic value-at-risk. It has a variational dual mirroring the entropic formula (an inverse temperature becomes a transport price), occupies a definite place in the risk hierarchy, and provably accounts for the reachable catastrophes the entropic measure ignores; we verify both dualities numerically. Driving the transport radius by belief entropy then yields a closed-form robust dynamic-programming operator whose caution contracts as the belief sharpens, with a certified safety sandwich and a sharp safety switch.
Abstract.
In this paper, we consider a noncollaborative game where each player faces two types of uncertainty: aleatoric uncertainty arising from inherent randomness of underlying data in its own decision-making problem and epistemic uncertainty arising from lack of knowledge and statistical information on the rivals’ risk preferences. By assuming that players are risk-averse against aleatoric uncertainty and risk-neutral on the epistemic uncertainty, we propose a Bayesian risk minimization model to describe players’ interactions. We investigate existence and uniqueness of a continuous monotone equilibrium arising from the game, termed CMBNRE, where “R” is used to emphasize the game is concerned with risk minimization as opposed to utility maximization in existing Bayesian Nash equilibrium (BNE) models. We derive sufficient and/or necessary conditions for ensuring existence of CMBNRE from two perspectives: monotonicity of a continuous BNRE and continuity of a monotone BNRE. Continuous BNE emphasizes continuity of each player’s response function with respect to variation of the type parameter whereas monotone BNE focuses on monotonicity of the player’s response function and/or the rival’s response functions. The former is derived by virtue of Schauder’s fixed point theorem and the latter is established by other fixed point theorems based on single-crossing and quasi-supermodularity; the proposed CMBNRE model synthesizes the two modeling frameworks. We discuss numerical methods for computing an approximate CMBNRE from a stochastic optimization perspective and apply the proposed model and computational schemes to a reinsurance competition problem. The test results provide some insights about how the indemnity parameter may affect a reinsurer’s behavior, market share, and profitability.
Unknown authors· SIAM Journal on Optimization· 4 citations· ⚡1
Inspired by Shapiro et al. [74], we consider a stochastic optimal control (SOC) and Markov decision process (MDP) under simultaneous epistemic and aleatoric uncertainties using Bayesian composite risk (BCR) measures. The proposed BCR-SOC/MDP model evaluates the risk of stagewise cost via a two-layer framework: the inner risk measure tackles aleatoric uncertainty conditional on a latent environment parameter, while the outer risk measure deals with the epistemic uncertainty of the inner risk under the Bayesian posterior. The resulting time-varying risk evaluation induced by Bayesian updating enables an information-adaptive risk-sensitive decision framework. Unlike [74], our policies are allowed to depend explicitly on the posterior belief, reflecting that accumulated information about epistemic uncertainty can influence the assessment of future aleatoric uncertainty and, consequently, the decision maker’s actions [79]. The new modeling paradigm subsumes several classical SOC/MDP formulations, including risk-averse and distributionally robust SOC/MDPs as well as partially observed and Bayes-adaptive MDPs, and generates so-called preference robust SOC/MDP models. Moreover, we derive conditions under which the BCR-SOC/MDP model is well-defined, show that finite-horizon BCR-SOC/MDP models can be solved via dynamic programming, and extend the analysis to the infinite-horizon case. Under standard conditions, we establish asymptotic convergence of the optimal values and optimal policies as data accumulate, and provide quantitative error bounds for several representative classes of risk measures. To enhance computational tractability, we develop a hyper-parameter discretization approach for the posterior belief space. Finally, we carry out numerical tests on a spread betting problem and an inventory control problem, demonstrating the effectiveness of the proposed model and numerical schemes.
The classic notion of strategyproofness implicitly assumes that a manipulating agent either possesses complete knowledge of what all other agents are going to report, or is willing to take the risk and act as if they know these reports. To capture the profound uncertainty of real-world voters, recent work introduced \emph{risk-avoiding truthfulness (RAT)} and the \emph{RAT-degree}, which quantifies the exact number of known reports required for a manipulation to be strictly safe. While the RAT-degree has been analyzed in settings such as single-winner elections, its implications for multi-winner voting remain unexplored. In this paper, we bridge this gap by extending the RAT-degree framework to approval-based committee (ABC) selection, focusing initially on the prominent Proportional Approval Voting (PAV) rule. We establish tight bounds on its susceptibility to safe subset manipulations, proving that PAV is immune to superset risk-avoiding manipulations given knowledge of at most $f = \lfloor \frac{n}{k+1} \rfloor - 1$ voters, but vulnerable when $f = \lceil \frac{n}{k} \rceil$. Recognizing that this degree of immunity may be insufficient in practice, we explore how to enhance strategic robustness by relaxing the proportionality requirement. We introduce a novel parameterized generalization of PAV, the family of $d$-RPAV rules, which encapsulates this inherent trade-off: a higher parameter $d$ yields stronger truthfulness and strategic robustness at the expense of weaker, relaxed proportionality guarantees. Specifically, we establish a generalized tight lower bound, proving that $d$-RPAV is completely immune to safe manipulation given knowledge of at most $f = \lfloor \frac{dn}{k+2d-1} \rfloor - 1$ voters.
The myopic escalation threshold is derived in closed form, characterise the optimal policy via dynamic programming, and it is proved that the optimal policy is a time-varying threshold with no shape assumption on the raw signal.
We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axioms and argumentation, but instead consequences, of higher-level modelling decisions. We show how epistemic and aleatoric uncertainty measures can be derived via decomposition of a subjective risk, based on a strictly proper loss. Reverse cross entropy provides a prominent example, where decomposition recovers the classic information-theoretic uncertainty terms. The same approach recovers numerous measures previously proposed across the UQ literature, providing them a common theoretical foundation. This suggests a new approach to UQ: given a modelling scenario and strictly proper loss, the corresponding epistemic and aleatoric terms are induced by the subjective-risk decomposition. We then extend our view to learning theory: we introduce and analyse subjective risk analogues of excess risk, approximation error and estimation error, and identify the connections to UQ. We consider this a first step towards a full learning-theoretic framework for uncertainty quantification.
R. Alamri, Michele Caprio, Gavin Brown· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.