Skip to content

A Noise-Robust Elicit-to-Optimize Framework for Distortion Riskmetrics via Inverse Reinforcement Learning

Jul 2026 · arXiv.org · Vol abs/2607.14373 · 0 citations · 27 references
Computer Science Economics

Abstract

We propose a noise-robust elicit-to-optimize framework that integrates inverse reinforcement learning (IRL) and reinforcement learning (RL) for eliciting agents'risk preferences and optimizing policies under a broad class of risk objectives characterized by distortion riskmetrics. On the elicitation side, we propose an adaptive Bayesian IRL method that infers agents'latent risk objectives from their noisy observed decisions, explicitly allowing agents to take stochastic and suboptimal actions. We establish the existence of a finite set of distinguishing questions that identifies the preferred distortion riskmetric within the candidate class and prove that the convergence rate of the algorithm is of order $O(\exp(-cm+O(\sqrt{m\log m})))$ under general settings, where $c>0$ is a constant and $m$ denotes the number of algorithm iterations. On the optimization side, we develop a model-free RL algorithm for optimizing policies under conditional distortion riskmetrics. By representing the objective as an integral of the conditional cost quantile function with respect to the distortion function, the method unifies distortion-riskmetric objectives. We optimize diverse risk objectives by extending the Proximal Policy Optimization (PPO) algorithm with policy, value, and quantile neural networks, where the quantile network estimates the full conditional cost quantile function and enables numerical evaluation of general risk objectives. A comprehensive empirical study demonstrates the framework's elicitation accuracy and effectiveness in complex financial environments.

View source

Similar papers

Preprint Aug 2026

Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making

RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty, and proves a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates.

D. Ganguly, Jan Křetinský · 0 citations
#artificial intelligence Preprint Aug 2026

Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk

An agent still learning its environment should be cautious while ignorant and bold once confident. The entropic value-at-risk captures this through a robust-optimization identity---a confidence level fixes the radius of a relative-entropy ball of alternative models---but that ball cannot reach catastrophes the nominal...

D. Ganguly, Jan Křetinský · 0 citations
Preprint Aug 2026

Finite-Time Analysis of Discounted Exponential-Utility Reinforcement Learning

This work establishes finite-time rates of $\tilde{O} (1/\sqrt{n})$ for the aforementioned two algorithms under asynchronous Markovian sampling, where $n$ is the iteration index and $\tilde{O}$ hides logarithmic expressions.

Ankur Naskar, A. VivekT, Aditya Kumar et al. · 0 citations
Preprint Aug 2026

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions

This work studies how many samples are necessary and sufficient to learn an $\varepsilon$-optimal robust policy under the average-reward criterion and achieves these rates using reduction-based plug-in procedures that select the reduction---nominal or robust---and its discount factor.

Yue-Peng Yang, Yuxin Chen, Yuejie Chi · 0 citations
Preprint Sep 2026

Reconciling Universal and Uniform Learning with $Q$-Aggregation

We study regression under bounded responses in terms of excess mean squared error. When the comparator class is finite, this setting is known as model selection aggregation, and achieving minimax excess risk requires improper learning algorithms. Contrary to this, in the universal learning framework no improperness is...

M. Høgsgaard, Patrick Rebeschini, Tobias Wegel · 0 citations
#data science Jul 2026

Simple-regret rates and minimax optimality of fixed-prior expected improvement in Matérn and squared-exponential RKHSs

We study expected improvement (EI) for minimizing a deterministic function $f$ in the RKHS $\mathcal H_k$ of a continuous positive-semidefinite kernel $k$ on a nonempty compact set $\mathcal X\subset\mathbb R^d$. Function values are observed exactly, and EI is computed from a fixed zero-mean Gaussian-process model with...

Emmanuel Vazquez, S. Petit · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.