Inverse Optimal Control (IOC) aims to infer the underlying cost functional of an agent from observations of its expert behavior. This paper studies the finite-horizon continuous-time inverse LQR problem from closed-loop state--input trajectories, where both the system matrices and the quadratic cost are unknown. The finite horizon induces a time-varying optimal gain, and this endogenous excitation serves as the structural mechanism that makes joint recovery possible. We quantify this mechanism through three computable conditioning indices, which measure state richness, gain-variation richness, and injectivity of a structured cost operator. Using these indices, we establish joint identifiability conditions for the inverse problem considered here. Crucially, these conditions guarantee recovery of the ground-truth system matrices $(A,B)$ and the true cost weighting matrices, rather than merely a behaviorally equivalent surrogate. We also develop a conditioning-aware sampled-data reconstruction method that reconstructs the gain $K(\cdot)$ and the closed-loop dynamics matrix $A_c(\cdot)$ from noisy measurements, recovers $(A,B)$ in closed form, and identifies the quadratic weights through a convex semidefinite program. We further establish the non-asymptotic perturbation bounds and the consistency of the full reconstruction method under sub-Gaussian observation noise, with explicit dependence on the same conditioning indices. Numerical experiments support the theory and illustrate the diagnostic value of the conditioning indices.
We study optimal input design over a finite horizon for linear dynamical systems. The goal is to minimize a weighted inverse-covariance (information) criterion subject to an energy budget. The set of covariances achievable by causal policies is convex but lacks a tractable explicit description, ruling out projection-based methods. We show that Frank--Wolfe applies naturally: each linear minimization subproblem is a budget-constrained finite-horizon linear quadratic (LQ) problem, solvable by a Riccati recursion and one-dimensional bisection over a Lagrange multiplier. Using smoothness of the objective over the feasible set, we establish an $\mathcal{O}(1/M)$ convergence rate for the objective value, while strong convexity yields an $\mathcal{O}(1/\sqrt{M})$ rate for the iterates. We further extend the framework to input design for system identification with unknown dynamics and adaptive online LQR, and illustrate the approach numerically.
Fethi Bencherki, Bruce D. Lee, Nikolai Matni et al.· 0 citations
We consider the problem of objective inference in the context of receding-horizon linear-quadratic regulator (LQR). In this setting, we are given sequential state-action observations, where each observed action is the first control of a newly solved finite-horizon LQR problem. We characterize when the objective of that problem is uniquely identifiable from these observations and when additional observations provide no new information about the objective. We then analyze action prediction at unseen states and show that all objectives reproducing the observed actions yield identical actions throughout the affine hull of the observed states; outside this hull, we derive an upper bound on the prediction error. % and deriving a prediction-error bound outside this hull. Additionally, we show that, when only the linear objective terms are unknown, exact prediction holds at every state. Finally, numerical results show that, even under stochastic observation noise, re-optimizing an inferred objective enables accurate action prediction at unseen states across different planning horizons.
Zhi-Yuan Jin, Jingqi Li, David Fridovich-Keil· 0 citations
This paper studies the inverse reinforcement learning (RL) problem for linear-quadratic mean-field (MF) social optimization. The considered system features multiplicative noise and indefinite cost weights, which violate standard convexity assumptions and pose analytical challenges. The goal is to recover unknown social cost weights from expert demonstrations and reproduce the optimal control policies. This requires solving coupled stochastic algebraic Riccati equations and Lyapunov equations with unknown system dynamics. To this end, we first propose a model-based inverse RL algorithm with two sequential loops that separately handle individual and MF dynamics, and we prove its convergence and closed-loop stabilizability. Moreover, we characterize the non-uniqueness of the recovered cost weights. To eliminate reliance on system dynamics, we develop a model-free inverse RL algorithm using integral RL and least-squares identification, which requires only measured trajectory data satisfying mild rank conditions. Finally, numerical simulations validate the effectiveness of the proposed approaches.
In adaptive control, parametric uncertainties in linear-in-parameter form consist of unknown parameters and known regressor signals. Convergence of the unknown parameters to their ideal values requires the regressor to satisfy a persistent excitation (PE) condition, which depends on future data and is therefore infeasible to guarantee online. Memory-based parameter update laws address this by enabling ideal parameter convergence under the online-verifiable finite excitation (FE) condition. In this paper, a new algorithm is proposed to construct a memory term via the Modified Gram-Schmidt orthogonalization procedure for a class of multi-input multi-output nonlinear systems with an unknown diagonal control effectiveness matrix and bounded nonparametric uncertainties. Under the finite excitation condition, the constructed memory term yields an identity coefficient matrix in the parameter estimation error dynamics. The identity coefficient matrix eliminates the need for time-varying adaptation gains, enables an explicit ultimate bound on the parameter estimation error, and preserves the structure of the nonparametric uncertainty bound under the memory term. Building on this, a combined adaptation law is developed for controller gain estimation under FE. The closed-loop tracking and estimation errors are shown to decay exponentially to a neighborhood of the origin, characterized by an explicit ultimate bound, with a decay rate that depends solely on user-defined gains and system constants, independent of the level of regressor excitation. This removes the dependence of the convergence rate on the level of regressor excitation, a key limitation of existing approaches such as concurrent learning, memory regressor extension, and DREM.
To address the problem of an unknown performance index in discrete-time (DT) linear systems with unobservable states, this paper investigates the output-feedback inverse reinforcement learning (IRL) problem and proposes a data-driven output-feedback off-policy inverse Q-learning algorithm. The proposed method relies solely on input–output trajectory data from an expert system. It does not require a system-dynamics model or state information to identify the unknown cost function and learn an optimal control policy. First, for cases where the system dynamics and the target control gain are known, we propose a model-based method. Building on this foundation, we further develop a model-free method that does not require a system-dynamics model or state information. The proposed algorithm uses state reconstruction to transform the state-feedback Q-function equation into an input–output form. Simulations using an F-16 aircraft model demonstrate that the learner system can approximate the expert trajectories and effectively reconstruct the performance-index function.
Zheng-Lin Li, Li-Qi Zhang, Xi-Sheng Dai et al.· Actuators· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.