Skip to content

Stochastic Stability of Nonlinear MPPI via Contraction Theory and Control Lyapunov Functions

Jul 2026 · arXiv.org · Vol abs/2607.06945 · 0 citations · 20 references
Computer Science Engineering Mathematics

TL;DR

It is shown that finite-sample MPPI inherits the nominal contraction when its sampling-based update approximates this reference policy with sufficient accuracy, and satisfies a finite-horizon, high-probability localized mean practical stability bound with residual floors due to MPPI approximation error, Gaussian process noise, and bad sampling events.

Abstract

Model Predictive Path Integral (MPPI) control is directly implementable on nonlinear systems because its online update requires only forward rollouts of the dynamics, not gradients, linearizations, or convex optimization. However, this algorithmic flexibility does not by itself provide a closed-loop stability certificate. This paper establishes such a certificate through a stability-inheritance argument. We assume that there exists a deterministic nonlinear MPC policy whose disturbance-free closed loop is certified by a Control Lyapunov Function terminal cost and a contraction metric, and we show that finite-sample MPPI inherits the nominal contraction when its sampling-based update approximates this reference policy with sufficient accuracy. The approximation error decomposes into a finite-temperature bias floor and a Monte Carlo term that vanishes at the inverse square-root rate in the sample count. Under an explicit small-gain condition, the resulting MPPI closed loop satisfies a finite-horizon, high-probability localized mean practical stability bound with residual floors due to MPPI approximation error, Gaussian process noise, and bad sampling events. The paper also gives an ISS-type restatement and a finite-horizon design procedure for choosing the localization set, temperature, and sample count.

View source

Similar papers

Preprint Aug 2026

Iterative State- and Control-Dependent Model Predictive Control: A Jacobian-Free Formulation for Constrained Nonlinear Systems

This paper presents an iterative model predictive control algorithm that stabilizes constrained nonlinear systems without evaluating a single plant derivative. By factoring the exact nonlinear dynamics into a pseudo-linear form using state- and control-dependent coefficients (SCDCs), we replace the standard nonconvex optimization with a sequence of constrained linear-quadratic programs. Refreezing the coefficient matrices along the previously predicted trajectory drives the iteration. Near the origin, we prove this sequence contracts to a unique fixed point. We explicitly bound the number of iterations required to reach any stopping tolerance, and we quantify the distance from the fixed point to a true Karush-Kuhn-Tucker point, showing this optimality gap vanishes quadratically as the state approaches the origin. Inflating the discrete algebraic Riccati equation generates terminal ingredients that guarantee recursive feasibility and asymptotic stability, even when the solver terminates early. We adapt the terminal penalty online, proving it remains uniformly bounded, and we secure output feedback through the block-observable canonical form, which extracts the exact system state directly from past inputs and outputs. Retaining the block-banded structure of the subproblem forces the computational cost to scale linearly with the horizon length $\ell$. This $O(\ell)$ complexity matches the iterative linear quadratic regulator (iLQR) but sharply undercuts the $O(\ell^3)$ scaling of dense sequential quadratic programming (SQP). Numerical studies on a saturated quadrotor, a nonholonomic integrator, and a nonminimum-phase plant illustrate the theoretical bounds and map how the algorithm compares with iLQR, SQP, and linear-parameter-varying MPC.

Mohammadreza Kamaldar · 1 citation
Preprint Aug 2026

Distributed model predictive control via finite-step control Lyapunov functions

As opposed to classical converse Lyapunov theorems, finite-step converse results are constructive and offer a different starting point: for sufficiently large finite step ahead, say $M$, in an explicit sense, any scaled norm can serve as a converse finite-step Lyapunov function. As for interconnected discrete-time systems, similar lines of argument lead to ``non-conservative''small-gain conditions. Motivated by this viewpoint, this paper develops a distributed model predictive control framework for constrained interconnected nonlinear discrete-time systems. Each subsystem solves one local optimization problem at each system time step instant by setting the local stage function in form of a local control finite-step like Lyapunov function, using time-aligned neighbor predictions, optimized state-constraint tightening radii, and a finite-step small-gain terminal inequality. Since received neighbor predictions need not equal the trajectories generated by future receding-horizon optimizations, the nominal converse certificates do not alone ensure recursive feasibility or stability. We therefore develop shift-compatible constraint margins, a local one-step terminal feasibility test for networks that are affine in control, and an analytical bound for the prediction and reoptimization mismatch. The resulting analysis gives recursive feasibility, constraint satisfaction, and a practical $M$-step Lyapunov estimate, with asymptotic convergence when the prediction and reoptimization mismatch bound tends to zero. For constrained linear networks, the conditions reduce to finite-dimensional matrix, QP, and SOCP tests. The framework is specialized to current sharing and terminal bus voltage safety under DC/DC power converters'operational constraints in a two-DGU DC microgrid evaluated on a small laboratory-scale prototype.

N. Noroozi, Maryam Sharifi · 0 citations
Preprint Sep 2026

Parallel Policy-Gradient Methods for Parameter Optimization of Nonlinear Feedback Controllers

Structured feedback controllers provide rigorous stability guarantees, but often require manual parameter tuning to achieve good closed-loop performance. Policy-gradient methods offer a systematic approach to parameter optimization; however, conventional gradient evaluation requires sequential forward state rollout and backward costate propagation. This letter develops a time-parallel policy-gradient framework for discrete-time nonlinear control-affine systems. We derive the policy-gradient expression where the state and costate rollouts required for policy-gradient evaluation are formulated as residual-minimization problems and solved using Gauss-Newton (GN) iterations with parallel associative scans. For closed-loop systems that are globally asymptotically stable and locally exponentially stable, we show that the residual-minimization problems satisfy a local Polyak-Lojasiewicz (PL) inequality and that the GN iterates converge locally at a quadratic rate. Moreover, the PL constant, the size of the convergence neighborhood, and the quadratic convergence bound are independent of the rollout horizon T. We also prove that, for any finite horizon T, the state solver recovers the exact trajectory from any initialization in at most T iterations. Finally, an inertia-wheel pendulum example with interconnection and damping assignment passivity-based control (IDA-PBC) demonstrates improved closed-loop performance and the computational benefits of the proposed parallel policy-gradient framework.

A. Nguyen, Leilei Cui · 0 citations
Preprint Aug 2026

Stabilizer Design for Policy Iteration in Stochastic Linear Quadratic Control: A Spectrum-Assignment Approach

Policy iteration (PI) is an important reinforcement learning tool for solving optimal control problems which includes an initialization stage, i.e., the search for an initial stabilizing controller. However, the initialization stage typically relies on complete model information, thereby imposing substantial constraints on the initialization of model-free PI. For stochastic systems with multiplicative noise dependent on state and control, the stability is not ensured by Hurwitz conditions as in the deterministic case, but rather by a Lyapunov-type inequality that incorporates both drift and diffusion terms. Therefore, the corresponding model-free PI initialization problem is more challenging. To this end, a novel spectrum assignment method is proposed to obtain an initial stabilizer for PI in continuous-time indefinite stochastic linear quadratic control. With the help of the Lyapunov-type operator's spectrum, the original system is gradually approximated from the stable auxiliary system by adjusting a cumulative factor, thereby obtaining a stabilizing control gain. Furthermore, by leveraging system data and adjusting the cumulative factor, we design a model-free algorithm that does not rely on an initial stabilizing policy and can achieve optimal control. Finally, simulation results are provided to validate the effectiveness of the proposed methods.

Xinyu Cao, Bing-Chang Wang, Ying Cao · 0 citations
Preprint Aug 2026

Policy Iteration for Linear-Quadratic Stochastic Differential Games with State- and Control-Dependent Noise

This paper presents a novel sequential policy iteration (PI) method for stochastic differential games with state- and control-dependent noise. The updates preserve mean-square stability, so that the iteration is well posed. We further derive a closed-form expression for the Fr\'echet derivative of the sequential PI map at a Nash equilibrium. The resulting characterization reveals how control-dependent noise, policy-evaluation sensitivity, and update ordering govern local error propagation, and yields explicit sufficient conditions for local linear convergence. Since finding an initial stabilizing solution is a major challenge in policy iteration, we also propose a homotopy-based initialization that ensures a valid starting point. The effectiveness of the proposed PI algorithm and the analytical results are verified through a numerical example.

Karl Handwerker, Felix Thömmes, Lucas Günther et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.