Skip to content

Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning

Sep 2026 · 0 citations · 30 references
Computer Science

TL;DR

This workCertified safety curation answers with a filter-then-clone pipeline: a state-only value trained from segment comparisons scores whole trajectories, Learn-then-Test calibration certifies a selection threshold under a distribution-free bound on the unsafe fraction of the selection, and behavior cloning follows.

Abstract

Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget. Certified safety curation answers with a filter-then-clone pipeline: a state-only value trained from segment comparisons scores whole trajectories, Learn-then-Test calibration certifies a selection threshold under a distribution-free $(\alpha, \delta)$ bound on the unsafe fraction of the selection, and behavior cloning follows. What is certified is the training set, not the policy, whose cost we report rather than bound. Where no threshold attains the target, the procedure refuses. We are not aware of prior work certifying the composition of a training set for offline RL or imitation. Oracle controls justify the design: reweighting individual transitions fails even with an exact value, so the value selects whole trajectories. The policies satisfy the cost budget on twelve of fifteen DSRL tasks, matching a clone of the ground-truth safe subset, which needs a label on every trajectory. Retrained on the certified selection, the strongest full-label method becomes safe where no setting of its own cost target rescues it. Refusal is predictable: the certificate's probability has a closed form in the purity of the pool's top quantile, which the calibration sample estimates and through which the scorer enters.

View source

Similar papers

Conference Open access Sep 2026

Persistent Safety Set Guided Offline Safe Reinforcement Learning

A framework for learning control barrier functions (CBFs) using a novel generalized Bellman operator is developed, yielding a persistent safety set from which the agent can remain safe indefinitely, and a new reward maximization algorithm is proposed that effectively exploits the learned persistent safety set for rewar...

A. Choudhury, J. Brahmanage, Akshat Kumar et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Evaluation Metrics for Safe Reinforcement Learning

Evaluation metrics for safe RL are introduced that address each of these concerns and in addition allow for aggregation across tasks and safety bounds and an open-source evaluation suite to support the reliable characterization of safety in future safe RL research is provided.

Lindsay Spoor, A. Plaat, T. Moerland · 0 citations
#machine learning Preprint Sep 2026

Learning Chance-Constrained MDPs with Bellman Distributional Certificates

Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is...

Chen-Bei Lu, Hong-Yu Yi · 0 citations
#artificial intelligence Preprint Oct 2026

FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training

Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget. Three coupled issues remain unresolved. A future-risk model trained from behavior trajectories need not estimate t...

Miao-Bo Hu, Shu-Hao Hu, Xiao-Bo Guo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

CANOPY (Coverage-ANchored On-PolicY RL), a minimalist protocol attacking both Signal starvation and policy drift, internalizes long-horizon capability directly into a small open model; the complete training stack is planned to be released at https://github.com/AlibabaResearch/SignalCoverageRL.

Li-Ming Pu, Xiao-Xiao Li, Yi-Fu Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs

This work develops a general, statistically efficient framework for last-iterate convergence in structured Constrained MDPs (CMDPs), and validate the stabilising effect predicted by the theory on a synthetic linear CMDP: the regularised method exhibits stable last-iterate behaviour, whereas its unregularised counterpar...

Nam Phuong Tran, Trinh Ha Mai Huynh, T. P. Le et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.