Skip to content
Preprint

Fundamental Limitations of Data-Driven Control: A Statistical Decision Perspective

Aug 2026 · 0 citations
Engineering Computer Science

TL;DR

This contribution develops a statistical decision framework for data-driven control, in which a controller is evaluated by its risk, defined as the expected performance degradation relative to the oracle model-based controller, and by its average risk over the parameter space.

Abstract

Substantial research efforts have been devoted to the design of data-driven controllers; however, comparatively less is known about their statistical performance and fundamental limitations. This contribution develops a statistical decision framework for data-driven control, in which a controller is evaluated by its risk, defined as the expected performance degradation relative to the oracle model-based controller, and by its average risk over the parameter space. Within this framework, we propose a collection of design principles for data-driven controllers. We further derive lower bounds on risks by combining the bias-variance decomposition with the Cram\'er-Rao inequality. In particular, the optimal bias that attains the lower bound for the average risk is determined by calculus of variations, thereby making the bias-variance tradeoff in data-driven control explicit. Moreover, the derived bound reveals a ``waterbed''effect in data-driven control: any improvement in risk relative to the lower bound over one region of the parameter space must be compensated by deterioration elsewhere. We illustrate the proposed framework on two canonical data-driven control problems: optimal feedforward control and the linear quadratic regulator benchmark. By comparing several representative data-driven controllers with the derived lower bounds, we sharpen the statistical interpretation of existing methods and reveal quantitative limitations that no controller design can avoid.

View source

Similar papers

Preprint Aug 2026

Achieving First-Order Statistical Improvements in Data-Driven Optimization: From No-Free-Lunch to Amplified Decision Perturbation

Recent proliferation of data-optimization integration has led to a range of methods that aim to improve the statistical performance of data-driven optimization decisions. However, while many of these methods are motivated intuitively from a robustness or regularization perspective, their resulting statistical benefits are often unclear and, even if available, are established on a case-by-case basis. We provide a systematic dissection of data-driven optimization formulations using the view of"directionally perturbed"empirical optimization (EO). Specifically, this umbrella of formulations, which we call"EO+", covers many existing data-driven optimization methods, including regularization, distributionally robust optimization, transfer learning, and analogous methods for contextual optimization. On the one hand, we argue that without additional, correctly specified, side information, any EO+ method can result in at most second-order improvements. This provides a negative conclusion, namely ``no free lunch is possible", on the statistical power of EO+. On the other hand, we show that when leveraging side information that is geometrically effective, achieving first-order improvements is possible by choosing hyperparameters that are significantly larger than what is typically suggested in the literature. Moreover, we construct a principled methodology based on excess risk estimation, via either system knowledge or bootstrap resampling, to maximize the first-order gain. We demonstrate how this gain connects to the control-variate principle, a variance reduction technique in the Monte Carlo simulation literature, which helps explain why geometrically effective side information is necessary.

Henry Lam, Tianyu Wang · 0 citations
Preprint Jul 2026

Statistical Inference for Scenario-Based Dynamic Optimization under Uncertainty

Motivated by batch and semi-batch process operation, we study finite-horizon open-loop dynamic optimization problems with uncertain parameters. A common computational approach replaces the expected performance criterion by an average over finitely many sampled parameter realizations. We develop a statistical theory for the resulting sample-based optimal value as an estimator of the population optimal value. The analysis is based on a stability estimate showing that terminal losses depend Lipschitz continuously on the time-integrated control, which records the cumulative input delivered up to each time. This estimate yields a functional central limit theorem for the sample-based objective and a statistical limit theorem for the corresponding optimal value error. As a consequence, we obtain confidence intervals for the population optimal value. When the population optimizer is unique, the limit is Gaussian and leads to a plug-in confidence interval. When multiple optimal policies may exist, we use a subsampling confidence interval that does not require uniqueness. The methodology is illustrated on two fed-batch case studies in which feed-rate profiles are optimized under parametric uncertainty.

Aurya Javeed, Johannes Milz · 0 citations
Review Aug 2026

On the Structural Limits of Machine Learning Decision Systems: An Information-Theoretic, Interaction-Based, and Stochastic-Dynamical Perspective

This work examines intrinsic limits of data-driven decision systems from an information-theoretic and interaction-based perspective and describes decision systems, including LLM-integrated agent architectures, as feedback-driven stochastic processes where state-dependent dynamics may induce emergent macroscopic behavior.

N. R. Barraza, G. Pena · 0 citations
Aug 2026

A Decay Rate for Steady Guaranteed In‐Control Performance in Statistical Process Control Models With Cautious Parameter Learning

Control charts with parameter estimates from small Phase I samples are at risk of providing inaccurate signals due to poor model fit. In turn, Guaranteed In‐Control Performance (GICP) approaches were introduced to account for the resulting variability in the conditional average run length. Combined GICP and Cautious Learning (GICP/CL) procedures were then proposed to mitigate the loss in sensitivity associated with GICP approaches. However, the performance of GICP/CL approaches is hitherto not fully explored. Previous research suggests that the convergence rate of the standard error, that is commonly used to adapt the control limits in GICP/CL frameworks, results in an unwanted gradual loss of detection power. This study explores the issue and shows that, when using the convergence rate of the standard error to adapt control limits, control charts calibrated for GICP have their sensitivity gradually decreased due to high variability in IC performance. Causes of high variability in the IC performance are small Phase I samples and low parameter updating frequencies. An alternative control limit adaption rate for steady GICP performance is proposed and recommendations for practical application are put forward.

Alexander Wendler, V. Tercero-Gómez, Dongping Du et al. · 0 citations
Preprint Jul 2026

When Persistency is not Exciting in Data-Driven Predictive Control

Understanding how to collect data that is meaningful for control purposes is of paramount importance in data-driven control. While existing approaches have primarily relied on the satisfaction of a rank condition to assess the quality of an experiment, we show that satisfying it is not always sufficient to achieve satisfactory closed-loop performance. Focusing on scenarios where white-noise-like excitation cannot be used for data collection, we examine the frequency-domain implications of linear behavioral representation. This analysis demonstrates that data must both satisfy the rank condition and excite the frequencies of interest for the control goal, thereby laying the foundations for control-oriented experiment design tailored to direct data-driven approaches. These findings are reflected in our numerical results. Data-enabled predictive controllers that rely on data satisfying the rank condition but neglect the tracking control goal result in a closed-loop system that cannot track the selected reference.

G. Giacomelli, Chuyu Lu, Siep Weiland et al. · 0 citations
Jul 2026

A subspace approach to data-driven predictive control for linear parameter-varying systems

This paper presents a subspace data-driven predictive control method for linear parameter-varying (LPV) systems. Starting from an affine LPV state-space model in innovation form, we derive a multi-step predictor that separates the effects of past data, future inputs, scheduling trajectories, and innovations. By projecting this representation onto the row span of lifted input-output-scheduling data, we obtain an asymptotically unbiased data-driven predictor that can be embedded directly in a receding-horizon control problem, without explicitly identifying an LPV model. To make the resulting LPV data-driven predictive control (DDPC) formulation tractable, we introduce an LPV extension of $\gamma$-DDPC based on an LQ factorization. This formulation fixes the number of online decision variables independently of the length of the dataset. A reduced-order predictor is then proposed to curb the exponential growth of scheduling-dependent regressors, which also relaxes the persistence-of-excitation condition. Simulation studies, including an unbalanced-disk example, show that the proposed controller achieves good tracking performance and, compared to existing LPV DDPC schemes, achieves better robustness to measurement noise and reduced computational cost, making multi-step LPV DDPC practically deployable, even with longer past horizons.

Federico Porcari, C. Verhoek, V. Breschi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.