Together, the results clarify how the input geometry shapes the kernel features and fundamentally impacts its generalization properties and clarify how the input geometry shapes the kernel features and fundamentally impacts its generalization properties.
Abstract
We study kernel ridge regression under anisotropic Gaussian data, where the input covariance decays as a power law with exponent $\alpha\geq 0$ for polynomial inner-product kernels. We derive asymptotically sharp expressions for the kernel spectrum and the generalization error in the polynomial high-dimensional regime $n=\Theta(d^\kappa)$, revealing how anisotropy reshapes the learning curves. For weak anisotropy ($0<\alpha<1$), the problem remains effectively high-dimensional and retains some features of the isotropic case, while departing from it in others: the variance still peaks at integer sample complexities $\kappa\in\mathbb{N}$, but these peaks are progressively damped as $\alpha$ grows; meanwhile, for targets strongly aligned with the data's principal directions, the bias drops at fractional sample complexities, decoupling the bias transitions from the interpolation peaks. For strong anisotropy ($\alpha>1$), the effective dimension of the problem is constant, and the variance stops depending on sample size altogether, plateauing under ridgeless interpolation or vanishing at an explicit rate under fixed ridge penalty. The bias undergoes a sharp transition governed by the target's decay rate: below a threshold, learning is abrupt rather than gradual; above it, the bias decays as a power law that recovers the classical source and capacity rates. We finally specialize these results to single-index targets, showing how the alignment of the index with the data's principal directions determines the effect of anisotropy on learning. Together, our results clarify how the input geometry shapes the kernel features and fundamentally impacts its generalization properties.
Kernel ridge regression is a standard method for functional data analysis, but its exact behavior is less understood. We study tensor-product kernel ridge regression for estimating the $r$-th moment function of a random function based on noisy discrete observations. The formulation includes mean estimation, covariance estimation, and higher-order moment estimation in a single framework. Our main result gives a precise $1+o_{\mathbb{P}}(1)$ expansion for the $L^2$ error at each admissible regularization parameter. The expansion consists of bias and three variance terms corresponding respectively to variation across the independent sample paths, latent signal variation at each sample point, and variation from measurement errors, identifying the refined error structure underlying functional data. As applications, we show that KRR attains the minimax rate for source smoothness $s \leq 2$ but becomes suboptimal in the sparse regime for $s>2$ due to saturation. A technical ingredient is a set of concentration inequalities for $U$-statistics suited to the dependent product structure of functional observations.
This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally. We consider a generalized spiked population covariance model with multiple latent factors, where the number of spiked eigenvalues may remain finite or increase with $n$, and the spiked eigenvalues may be bounded or diverge at arbitrary rates. Beyond characterizing the impact of covariance spectra, we reveal a new mechanism underlying benign overfitting: the prediction behavior of ridgeless interpolation is fundamentally governed by the alignment between the regression coefficient $\boldsymbol\beta$ and the spiked eigenspaces of the population covariance matrix. In particular, we show that the signal energy distributed along latent spike directions determines whether interpolation leads to benign, tempered, or catastrophic overfitting. Our theoretical framework establishes sharp prediction risk limits under minimal moment conditions, requiring only finite fourth moments rather than Gaussianity. We characterize how the number, strength, and geometric structure of the spikes jointly influence the double-descent phenomenon. These results provide a unified understanding of when latent covariance structures facilitate or hinder generalization in overparameterized regression.
We investigate the geometric fluctuations of principal subspaces for high-dimensional covariance matrices through the squared Frobenius $\sin\Theta$ distance between the sample and population eigenspaces associated with the $r_p$ largest eigenvalues. An explicit first-order expansion and a central limit theorem are established for this subspace distance. The theory allows the subspace dimension to diverge subject to $r_p=o(n)$, where $n$ is the sample size. It also permits a diverging spectral norm of the population covariance matrix, population spikes of different orders, and repeated or closely spaced spikes. This sharp characterisation captures features of the subspace estimation error that are not reflected in existing perturbation bounds. As applications, we derive an explicit asymptotic expansion for the expected PCA excess risk and a refined error bound for distributed PCA. In both cases, existing upper bounds can increase with the spiked-block condition number when some leading spikes become stronger, whereas our results show that the corresponding estimation errors need not increase and may instead decrease. Numerical experiments reproduce this contrasting behaviour and demonstrate the finite-sample accuracy of our theoretical findings.
We show that, up to isotropic scaling, the Gaussian RBF reproducing kernel Hilbert space (RKHS) is asymptotically isometric to Euclidean space in the large bandwidth limit. This strongly suggests that kernel-based constructions reliant on metric properties of the RKHS will yield results for Gaussian RBF kernels that similarly approach those of linear kernels for large bandwidths. The asymptotic behavior of Gaussian CKA can be understood in this light. We further consider kernel PCA, showing that Gaussian RBF eigenvalues, eigenprojections, and principal components all converge to those of classical (linear) PCA as bandwidth $\sigma \rightarrow \infty$. For a given data representation, both the RKHS feature embeddings and the orthogonal PCA eigenframes of the two kernel types differ asymptotically by a geometric similarity transformation, up to a residual of size $O \left (\frac{\rho}{\sigma} \right )^2$, where $\rho$ is a measure of geometric eccentricity of the representation, equal to the ratio of maximum to median pairwise distance between data examples. Experiments over a diverse collection of data sets demonstrate that $\rho$ provides a simple and reliable predictor of dataset-specific convergence behavior in the top principal directions.
Theoretical guarantees for kernel regression are typically formulated in terms of smoothness, but practical accuracy depends critically on how the design resolution compares with the target and kernel lengthscales. We develop a finite-sample theory for Mat\'ern regression on periodic domains with quasi-uniform designs, covering noiseless interpolation and noisy kernel ridge regression. Our minimax result shows that accurate recovery requires a design dense enough to resolve the target lengthscale and sufficient information at that scale to overcome noise. We prove that Mat\'ern interpolation additionally requires the design to resolve the kernel lengthscale: if the kernel is too short relative to the point spacing, an additional error remains even when the target is well resolved. For noisy Mat\'ern regression, we derive a three-term fixed-ridge risk characterization consisting of target-scale bias, kernel-scale bias, and variance, and show that optimizing over the ridge parameter yields four distinct contributions. When the target is no more than twice as smooth as the kernel, choosing a kernel lengthscale longer than the target does not worsen the oracle risk, whereas choosing one too short can. Thus these resolution conditions determine when accurate recovery becomes possible, while smoothness determines how rapidly the error decreases thereafter.
We study the asymptotic spectral properties of high-dimensional Spearman correlation matrices for scale-mixture data. We consider observations of the form $x_t=\sigma_t \xi_t \in \mathbb{R}^N,$ where the coordinates of $\xi_t$ are i.i.d.\ and the scalar mixture variable $\sigma_t$ is shared by all coordinates. Under natural symmetry assumptions, the coordinates of $x_t$ are pairwise uncorrelated in both the Pearson and Spearman sense. Nevertheless, they are not independent when the mixture variable is non-degenerate. We show that this higher-order dependence survives the rank transformation and leaves a nontrivial spectral signature. In the proportional regime $N/T\to q\in(0,\infty),$ the empirical spectral distribution of the Spearman correlation matrix converges almost surely to a generalized Mar\v{c}enko--Pastur law governed by the limiting distribution of an effective rank variance. We also formulate a broader latent-variable extension, which covers, in particular, some scale-mixture models with correlated directional components. We discuss solvable examples and numerical approximations, motivated in part by heavy-tailed data in robust multivariate statistics, econometrics, and finance.
J. Bouchaud, Pierre Bousseyroux, Tomas Espana et al.· 0 citations
Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.
What does it take to trust AI-driven HVAC optimization? Our AI Model Factory combines agents, machine learning, reinforcement learning and deterministic checks in a governed workflow designed for messy, real-world building data. The post We built an AI factory for HVAC control appeared first on GPT-Lab.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.