Skip to content
Preprint

Physics-informed Reinforcement Learning for Stochastic Reach-Avoid Analysis

Aug 2026 · 0 citations · 42 references
Engineering Computer Science

TL;DR

A physics-informed RL (PIRL) framework that combines the complementary strengths of PINNs and RL for stochastic reach-avoid analysis, and develops a scheduled PIRL algorithm in which temporal-difference actor-critic learning first guides the critic toward a meaningful approximation of the reach-avoid value function.

Abstract

Stochastic reach-avoid analysis of controlled dynamical systems is an important tool for safety-critical control under uncertainty, in which the reach-avoid probability is characterized by a Hamilton-Jacobi partial differential equation (PDE). However, solving this PDE using conventional numerical methods becomes computationally intractable as the system dimension increases. Physics-informed neural networks (PINNs) may converge to inaccurate local minima when trained primarily through PDE-residual minimization. Reinforcement learning (RL) offers a scalable alternative, but its learned value functions may be inaccurate or inconsistent with the governing PDE. This paper proposes a physics-informed RL (PIRL) framework that combines the complementary strengths of PINNs and RL for stochastic reach-avoid analysis. We develop a scheduled PIRL algorithm in which temporal-difference actor-critic learning first guides the critic toward a meaningful approximation of the reach-avoid value function. PDE-residual and boundary-condition losses are then introduced progressively to enforce consistency with the governing PDE and its boundary conditions. The proposed method mitigates the failure modes of conventional PINN techniques while achieving accuracy comparable to that of successfully trained PINNs. The effectiveness of the proposed framework is demonstrated through two case studies.

View source

Similar papers

#machine learning Preprint Sep 2026

Learning to Solve Stochastic Controls with Unknown Drifts and Running Rewards: Theory, Algorithms and Convergence

We study continuous-time and possibly high-dimensional stochastic control problems where drift coefficients and running reward functions are unknown. Due to these missing model primitives, we take the exploratory, reinforcement learning (RL) framework of Wang, Zariphopoulou, and Zhou(2020) with relaxed controls and entropy regularization. The objective is to develop theoretically grounded, efficient and scalable RL algorithms to learn both the optimal value functions (which also solve the exploratory HJB equation) and optimal exploratory feedback control policies. When the diffusion coefficients do not contain control, we employ probabilistic representations of both the optimal value function and its gradient based on an auxiliary state process depending only on the diffusion part of the original dynamics. With a delicate analysis on some properly defined mappings and their fixed points, this leads to the introduction of our policy iteration algorithms and their convergence. We demonstrate the performance of our algorithms through various numerical examples. Finally, we study a special control-dependent diffusion case where probability representation of the Hessian is called for.

Jing-Sheng Ma, Gao-Zhan Wang, Jian-Feng Zhang et al. · 0 citations
Review Open access Aug 2026

Mathematical Foundations of Reinforcement Learning and Stochastic Control Systems

Reinforcement learning (RL) and stochastic optimal control both address a single underlying mathematical problem: how an agent should select actions over time, under uncertainty, to optimize a cumulative reward or cost criterion. This paper reviews the shared mathematical scaffolding uniting these two fields, tracing the progression from Bellman's dynamic programming and the Markov decision process (MDP) formalism through stochastic approximation theory, temporal-difference learning, policy-gradient methods, actor-critic architectures, and the Hamilton–Jacobi–Bellman (HJB) equation governing continuous-time stochastic control. Particular attention is given to the convergence-theoretic results that justify RL algorithms as legitimate stochastic approximation procedures: Robbins and Monro's foundational stochastic approximation method, Jaakkola, Jordan, and Singh's convergence proof for stochastic iterative dynamic programming, Tsitsiklis and Van Roy's analysis of temporal-difference learning with linear function approximation, and the policy-gradient theorem of Sutton, McAllester, Singh, and Mansour. The paper further examines deterministic policy-gradient methods, deep reinforcement learning's departure from classical convergence guarantees, and the connection between the discrete-time Bellman equation and its continuous-time HJB counterpart via viscosity solution theory. Comparative tables map core mathematical structures onto their representation assumption and guarantee type, contrast convergence rigor across tabular, linear, and nonlinear function-approximation regimes, and set the discrete-time RL and continuous-time control literatures against their shared and divergent mathematical machinery. The paper concludes that convergence guarantees degrade in a predictable, representation-dependent order as function approximation grows more expressive, and identifies extending stochastic approximation theory to nonlinear function approximation as the central future research prospect.

Shirish Prabhakarrao Kulkarni, C. Ashwini, M. Buvanasankari et al. · 0 citations
Preprint Aug 2026

Stable Multi-Step Rollouts via Uncertainty-Guided Hybrid Dynamics

A model-agnostic hybrid dynamics framework that blends a provably contracting nominal model with a flexible excursion model through an uncertainty-guided switching law is proposed, ensuring that each model operates within its reliability regime.

A. Maalberg, A. Neumann, J. Knobloch · 0 citations
Preprint Sep 2026

Gradient-Free Neural Hamilton-Jacobi Reachability for Scalable Safety-Critical Control

Hamilton-Jacobi (HJ) reachability provides a principled framework for synthesizing safety certificates and robust controllers for safety-critical robotic systems. However, applying reachability analysis to high-dimensional nonlinear systems remains challenging: classical grid-based solvers suffer from the curse of dimensionality, continuous-time neural solvers require accurate spatial value gradients, and reinforcement-learning-based approaches often suffer from weak boundary anchoring and non-stationary adversarial policy optimization. We propose a discrete-time neural reachability framework for control-disturbance-affine systems that learns backward reachable tubes (BRTs) and backward reach-avoid tubes (BRATs) through Bellman-Isaacs value propagation. Our key idea is to combine equation-driven self-supervision with structured policy learning: rather than computing explicit PDE-gradients, we exploit the bang-bang structure of optimal safety interventions to construct approximate teacher actions from gradient-free value probes, converting adversarial actor learning into supervised policy learning. To stabilize long-horizon value propagation, we leverage the learned actor to train the value function backward from the terminal boundary using a windowed temporal curriculum, where each window is used as the boundary condition for the next window. Across benchmark problems up to 80 dimensions, our method learns accurate reachability value functions while improving stability over existing learning-based solvers. We further demonstrate observation-space scalability on F1-tenth racing with over 16,000-dimensional egocentric inputs. The learned safety filter generalizes zero-shot to unseen tracks and transfers to a physical RC car, achieving real-time robust collision avoidance.

Ze-Yuan Feng, Ali Fuat Sahin, Santiago Thorup et al. · 0 citations
Preprint Aug 2026

Analytic Planning under Uncertainty with Moment Closure

Effective model-based reinforcement learning in stochastic environments requires planning that accounts for predictive uncertainty. Propagating full state distributions analytically offers a principled way to do this, but has traditionally required restrictive policy or reward structures to remain tractable. Consequently, modern deep reinforcement learning has largely retreated to either stochastic sampling, which introduces significant target variance, or deterministic point estimates that ignore predictive covariance entirely. We investigate whether distribution-aware planning is possible without these constraints. Using a quadratic action-value parameterization, we first reduce the Bellman backup to an expectation over the state-value function alone; the key idea is then a compatibility principle between the predictive transition distribution and the value function class, under which this expectation is analytic in the distribution's moments. We instantiate this principle with a Gaussian transition model paired with a radial-basis value function, yielding a closed-form backup that propagates both predictive mean and covariance. Empirically, our approach reduces target variance and yields well-calibrated predictive uncertainty under stochastic observations in continuous control, providing a principled framework for planning with learned distribution models.

Shishir Sharma, D. Precup · 0 citations
Preprint Jul 2026

Tikhonov-Regularized Physics-Informed Neural Networks for Terminal-State Distributed Optimal Control of Parabolic Partial Differential Equations

A regularized PINNs framework that incorporates Tikhonov regularization to solve terminal-state tracking optimal control constrained by parabolic partial differential equations is proposed, establishing a consistency result showing that PINNs minimizers nearly attain the continuous regularized objective under residual and quadrature approximation assumptions.

Q. Nguyen, T. Mai · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.