Skip to content

Long-Term Sequential Decision Making under Risk

Jul 2026 · arXiv.org · Vol abs/2607.19914 · 0 citations · 42 references
Computer Science

TL;DR

ERQDP is proposed, an enumeration-free and sampling-free method that solves a rank--quantile surrogate via exact DP (Dynamic Programming), evaluates candidate policies exactly by DP over return Probability Mass Functions (PMFs) on a discretized return grid (with an explicit rounding bound), and refines the surrogate in an anytime loop.

Abstract

We study finite-horizon MDP planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distribution and generally break Bellman optimality, so direct optimization by scenario-tree enumeration is intractable. We propose \textbf{ERQDP}, an enumeration-free and sampling-free method that solves a rank--quantile surrogate via exact DP (Dynamic Programming), evaluates candidate policies exactly by DP over return Probability Mass Functions (PMFs) on a discretized return grid (with an explicit rounding bound), and refines the surrogate in an anytime loop that reports an explicit upper--lower gap (certificate) for the target objective up to discretization budgets. Across tested benchmarks, ERQDP returns certified solutions or explicit residual gaps, enables fast risk-parameter sweeps with substantial runtime gains, and supports both risk-averse and risk-seeking behaviors.

View source

Similar papers

Preprint Aug 2026

Multi-period Value-at-Risk Constrained Portfolio Optimization via DC Programming

We study a multi-period portfolio optimization problem with finite-scenario Value-at-Risk (VaR) constraints, transaction costs, and diversification regularization. Using a finite-scenario VaR--CVaR identity, we derive a penalized difference-of-convex (DC) formulation over the underlying convex portfolio set. To solve t...

Thi Thu Van Nguyen · 0 citations
#machine learning Review Sep 2026

Tail-Influence Sampling for CVaR Policy Evaluation

Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to es...

P. Bourigault, Xiao-Tong Ji, Matthieu Zimmer et al. · 0 citations
Preprint Sep 2026

Learning to Solve Two-Stage Stochastic Unit Commitment Problems with Quality Guarantees

Two-stage stochastic Mixed-Integer Linear Programs are a canonical modeling tool to optimize power system operations under uncertainty, yet their extensive-form counterparts scale linearly with the number of scenarios and quickly become computationally prohibitive under day-ahead time constraints. We propose an Input C...

Andrea Fusco, Andrea Lodi, Lavanya Marla · 0 citations
#machine learning Preprint Sep 2026

Proximal Residual Value Functions for Consistent Planning and Real-Time Execution

We study two-timescale decision systems in which a planning layer periodically supplies a continuation-value function to a real-time optimizer that allocates arriving resources, with inventory placement as our motivating application. We propose an end-to-end reinforcement learning (RL) method for learning this function...

Harrison Waldon, Carson Eisenach, Akhil Bagaria et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.