On-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data are developed, specifically developing on-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data.
Abstract
This paper addresses infinite-horizon continuous-time stochastic linear quadratic optimal control problems with regime switching. We propose a paradigm shift from model-based design by adopting an adaptive dynamic programming approach, specifically developing on-policy and off-policy Q-learning algorithms that learn the optimal controller solely from online state trajectory data. The theoretical core of our work consists of a complete proof of the equivalence between the on- and off-policy architectures, alongside a rigorous analysis establishing the stability of the closed-loop system and the convergence of the algorithms to the optimal solution. For computational tractability, we implement these algorithms using vectorization and Kronecker product algebra. The theoretical results are corroborated by numerical case studies that clearly demonstrate the operational effectiveness and practical feasibility of the proposed model-free control strategy.
Linear Quadratic Stochastic Optimal Control (LQ-SOC) establishes a fundamental framework for steering noisy dynamical systems and has recently gained renewed interest in the machine learning community. However, current state-of-the-art policy-based methods suffer from prohibitive computational costs and instability due to their heavy reliance on full-trajectory simulation. To overcome these limitations, we propose a paradigm shift toward a value-based approach by revisiting Path Integral Control (PIC). Although standard PIC suffers from the same high-variance bottleneck as policy-based methods, we discover that by truncating and marginalizing the original path integral formulation, we can derive a temporal recursive form of the value function. Building upon this theoretical foundation, we propose the Path Integral Value Matching (PI-VM) algorithm. Specifically, we employ temporal-difference learning to approximate the recursive value dynamics, and further integrate the Girsanov theorem with experience replay to enable off-policy training. We benchmark PI-VM against SOTA policy-based methods across various SOC benchmarks and sampling tasks. Empirical results demonstrate that PI-VM matches SOTA precision with an order-of-magnitude efficiency gain in low-dimensional settings, while effectively mitigating mode collapse in high-dimensional scenarios. Consequently, PI-VM offers a scalable solution for solving complex SOC problems.
Bangyan Liao, Chenglei Yu, Yuchen Yang et al.· 0 citations
This paper presents a novel model-free algorithm for the finite-population linear quadratic Gaussian (LQG) decentralized social control problem with multiplicative noise. The state and control weights in the cost functional are not limited to be positive semidefinite. For both finite-horizon and infinite-horizon cases, the goal is to obtain a social optimum by solving two algebraic Riccati equations (AREs), without requiring prior knowledge of the system matrices. Then, we complete the design of a model-free algorithm for solving the decentralized social control problem. Especially, in the infinite-horizon case, the algorithm's convergence is based on analyzing the spectral property of the Lyapunov-type operator. The differences of reinforcement learning (RL) solutions between the finite-horizon and infinite-horizon cases are compared. Finally, the effectiveness of the proposed algorithm is demonstrated by a numerical example.
In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints. By means of an appropriate state augmentation, we reformulate the original problem as a constrained Markov decision process, in which both the cost and the constraint function exhibit an additive structure. We then prove that this formulation enjoys strong duality, thereby enabling us to reformulate the problem as an equivalent unconstrained one in the Lagrange dual framework. We propose a dual-ascent algorithm to solve the resulting problem and show that it converges to a deterministic Markov policy defined over the augmented state space that is both optimal and feasible. To accommodate continuous state-input spaces, we propose a dedicated learning algorithm to approximate the value function in an offline training setting, thereby significantly reducing the computational complexity of the online control phase. We then test our approach on a numerical example and demonstrate its effectiveness compared to online predictive control methods in terms of performance and computational complexity.
Francesco Cordiano, Kanghui He, B. de Schutter· 0 citations
This paper presents a novel sequential policy iteration (PI) method for stochastic differential games with state- and control-dependent noise. The updates preserve mean-square stability, so that the iteration is well posed. We further derive a closed-form expression for the Fr\'echet derivative of the sequential PI map at a Nash equilibrium. The resulting characterization reveals how control-dependent noise, policy-evaluation sensitivity, and update ordering govern local error propagation, and yields explicit sufficient conditions for local linear convergence. Since finding an initial stabilizing solution is a major challenge in policy iteration, we also propose a homotopy-based initialization that ensures a valid starting point. The effectiveness of the proposed PI algorithm and the analytical results are verified through a numerical example.
Karl Handwerker, Felix Thömmes, Lucas Günther et al.· 1 citation
Policy iteration (PI) is an important reinforcement learning tool for solving optimal control problems which includes an initialization stage, i.e., the search for an initial stabilizing controller. However, the initialization stage typically relies on complete model information, thereby imposing substantial constraints on the initialization of model-free PI. For stochastic systems with multiplicative noise dependent on state and control, the stability is not ensured by Hurwitz conditions as in the deterministic case, but rather by a Lyapunov-type inequality that incorporates both drift and diffusion terms. Therefore, the corresponding model-free PI initialization problem is more challenging. To this end, a novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control. With the help of the Lyapunov-type operator's spectrum, the original system is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtaining a stabilizing control gain. Furthermore, by leveraging system data and adjusting the cumulative factor, we design a model-free algorithm that does not rely on an initial stabilizing policy and can achieve optimal control. Finally, simulation results are provided to validate the effectiveness of the proposed methods.
Xinyu Cao, Bing-Chang Wang, Ying Cao· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.