This work proposes a Gaussian refit for kernel ridge regression by Anderson's inequality, which requires no moment assumptions and is calibrated at any confidence level via order statistics, and extends empirically to nonlinear constrained estimators and real spatial data.
Abstract
Assessing a single model fit requires a computable upper confidence bound for the gap between the fit and the unknown truth, as mean estimates ignore realization variance. Standard cross-validation margins are bottlenecked at order $n^{-1/2}$ by noise fluctuations, even when the true error shrinks faster. While wild refitting cancels this noise level, existing Rademacher sign methods degenerate for kernel ridge regression and rely on unobservable quantities. We propose a Gaussian refit for kernel ridge regression. By Anderson's inequality, the fit movement is monotone in the noise sizes, yielding a computable tail bound. Assuming only symmetric noise, the bound requires no moment assumptions and is calibrated at any confidence level via order statistics. Theoretically, using a worst-case envelope, the bound contracts at the minimax rate $O_P(n^{-2s/(2s+1)})$, correctly matching the prediction error. Empirically, using a practical data-driven envelope, the bound maintains full coverage within twice the true $95\%$ error quantile. By contrast, cross-validation exceeds this quantile by factors up to $51$, and by hundreds under infinite-variance noise. The procedure extends empirically to nonlinear constrained estimators and real spatial data.
Kernel ridge regression is a standard method for functional data analysis, but its exact behavior is less understood. We study tensor-product kernel ridge regression for estimating the $r$-th moment function of a random function based on noisy discrete observations. The formulation includes mean estimation, covariance estimation, and higher-order moment estimation in a single framework. Our main result gives a precise $1+o_{\mathbb{P}}(1)$ expansion for the $L^2$ error at each admissible regularization parameter. The expansion consists of bias and three variance terms corresponding respectively to variation across the independent sample paths, latent signal variation at each sample point, and variation from measurement errors, identifying the refined error structure underlying functional data. As applications, we show that KRR attains the minimax rate for source smoothness $s \leq 2$ but becomes suboptimal in the sparse regime for $s>2$ due to saturation. A technical ingredient is a set of concentration inequalities for $U$-statistics suited to the dependent product structure of functional observations.
This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally. We consider a generalized spiked population covariance model with multiple latent factors, where the number of spiked eigenvalues may remain finite or increase with $n$, and the spiked eigenvalues may be bounded or diverge at arbitrary rates. Beyond characterizing the impact of covariance spectra, we reveal a new mechanism underlying benign overfitting: the prediction behavior of ridgeless interpolation is fundamentally governed by the alignment between the regression coefficient $\boldsymbol\beta$ and the spiked eigenspaces of the population covariance matrix. In particular, we show that the signal energy distributed along latent spike directions determines whether interpolation leads to benign, tempered, or catastrophic overfitting. Our theoretical framework establishes sharp prediction risk limits under minimal moment conditions, requiring only finite fourth moments rather than Gaussianity. We characterize how the number, strength, and geometric structure of the spikes jointly influence the double-descent phenomenon. These results provide a unified understanding of when latent covariance structures facilitate or hinder generalization in overparameterized regression.
We introduce a flexible model for covariate-dependent multiple testing which can be encoded using a nonparametric Gaussian mixture model. Weight-localized predictive recursion (PRx), a new development in the methodology of Newton's predictive recursion algorithm, is then leveraged to estimate the components of this mixture model, allowing for recovery of the covariate-localized false discovery rate $\text{Pr}(H_i = 0|z_i,x_i)$ using a single, unified algorithm. This quantity represents the most direct extension of Efron's local false discovery rate to the covariate-dependent setting, and admits provable Bayesian FDR control properties under simple rejection rules. We introduce several procedures for estimating and thresholding the local false discovery rate, and show using various simulations and a real-data example that our procedures lead to increased power, tighter Bayesian FDR control, and more interpretable rejections. We furthermore show that this holds for fixed and randomized hypothesis labels, indicating that our proposed methods perform well under both frequentist and Bayesian interpretations of multiple testing.
For decades, the bootstrap has been a default tool for statistical inference because of its broad applicability and minimal analytic requirements. Although its validity is well understood for smooth parametric estimators, its theoretical properties for many modern semiparametric and machine-learning estimators remain largely unstudied. Nevertheless, bootstrap procedures are often used routinely in such settings, even when their validity is unknown and their computational cost is substantial. We develop the $V$-fold jackknife as a computationally efficient and theoretically justified alternative for semiparametric inference. It requires only $V$ leave-fold-out refits and uses the empirical dispersion of jackknife pseudo-values to quantify uncertainty, without deriving or evaluating an influence function. For regular asymptotically linear estimators of pathwise differentiable parameters, we show that, for fixed $V$, the Studentized $V$-fold jackknife statistic converges to a $t$-distribution with $V-1$ degrees of freedom, giving valid confidence intervals even though the jackknife variance estimator does not converge in probability. When $V\to\infty$, we establish consistency of the variance estimator at rate $V^{-1/2}$, allowing $V$ to diverge slowly, for example at rate $\log n$. We also develop simultaneous confidence bands based on the correct componentwise-Studentized limiting distribution. Finally, we extend the theory to generalized asymptotically linear estimators with diverging influence-function variance and slower-than-$\sqrt n$ convergence; scale invariance of Studentization eliminates the need to know the effective convergence rate. Simulations on the average treatment effect, Kaplan--Meier survival curve, and highly adaptive lasso dose-response curves confirm reliable inference, including where influence-function-based standard errors are anti-conservative or unstable.
Yi Li, Ashkan Ertefaie, M. J. van der Laan· 0 citations
A neural network is trained on simulated datasets drawn from a prior over a distribution family, using single independent draws of the root T_n - T(F) scored by the pinball loss, a proper scoring rule whose population minimizer is the posterior-predictive law of the root.
We study a class of Lasso based estimators obtained by applying a quadratic correction on the Lasso equicorrelation set. The penalty matrix determines both the magnitude and geometry of the correction and contains, among other cases, the isotropic Lasso--Ridge correction, least squares refitting, Gram proportional interpolation between the Lasso and least squares, and coordinate specific penalties. We first derive a closed form representation and isolate the positive gain component of the resulting prediction improvement. We then control the remaining stochastic linear term in expectation by localizing the random signed equicorrelation model around a deterministic reference support. This yields a finite sample expectation bound that explicitly accounts for the randomness induced by Lasso model selection. The resulting decomposition provides a unified framework for understanding when Lasso based quadratic corrections can improve prediction.
Guo Liu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.