Skip to content

Sharp Rates and a One-Line Correction for Spectral Representation Learning

Sep 2026 · 0 citations · 35 references
Computer Science Mathematics

Abstract

A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named at training time; the practitioner's question is when the off-the-shelf features are good enough and when they need fixing. Canonical correlation analysis, HGR maximal correlation, and the population optimum of the spectral contrastive loss all return the top-$k$ singular subspace of a cross-view dependence operator, justified by isotropy: if the task prior has no directional preference, that subspace is universally optimal. We show isotropy is the wrong hypothesis. The prior enters the transfer risk only through the task covariance $\Lambda=\mathbb{E}[\Delta\Delta^\top]$, and only through its compression onto the operator's leading singular directions; what matters is not whether $\Lambda$ is isotropic but whether its preferred directions are ordered consistently with the operator's spectrum. We prove matching two-sided rates---worst-case regret is exactly $1-1/\kappa(\Lambda)$, refines to $1-A_k$ for an alignment coefficient $A_k$, localizes to the top-$2k$ subspace, becomes second order under a spectral gap, and is improvable by no task-agnostic representation---and show why alignment is generic: incoherent preferences cancel in high dimension, and $T$ diverse tasks force $\alpha=\widetilde O(\sqrt{d_x/T})$, a quantitative account of why task diversity, not symmetry, makes self-supervised features transfer. The governing statistics cost $O(kd_x^2)$, and when they signal misalignment a one-line reweighting of the positive-pair term provably restores exact optimality. The result is a diagnostic that answers the practitioner's question from a small labelled budget and refuses when the task bank cannot support the width requested; on controlled data it takes a regret of $0.86$ down to $0.003$, and on a CIFAR-100 encoder it correctly predicts that no correction is needed.

View source

Similar papers

Preprint Aug 2026

Bandable Cumulant Tensors: Optimal Estimation and Applications in Non-Gaussian Data Modeling

Higher-order cumulants capture the non-Gaussian dependence that covariance misses, but they are hard to use in high dimensions. An order-$d$ cumulant tensor has $p^d$ entries, and the plug-in sample cumulant is generally not even rate-optimal under the tensor spectral norm. For ordered data, both difficulties admit one...

Run-Shi Tang, An-Ru R. Zhang, Yuefeng Han et al. · 0 citations
Preprint Sep 2026

Gradient-Only Online Convex Optimization with a Single Quadratic Leader

We study online convex optimization on a bounded Euclidean domain when the learner receives only one subgradient at its prediction and knows neither the horizon, a gradient bound, nor the losses'strong-convexity parameters. We present a scale-invariant algorithm that couples projected adaptive gradient descent to one c...

Kamiar Asgari · 0 citations
Preprint Aug 2026

Inference and Uncertainty Quantification for Streaming $r$-PCA

We address two open questions in streaming PCA via Oja's algorithm: sharp operator-norm convergence for general rank under sub-Gaussian data, and distributional inference for the resulting subspace estimator. Existing convergence analyses, even in the rank-one case, either assume bounded data or leave non-vanishing rem...

Hao-Shu Xu, Hong-Zhe Li · 0 citations
#machine learning Preprint Sep 2026

Restricted Eigenvalues Beyond Gaussian Width: Threshold Occupancy under Heavy Tails

Restricted eigenvalue (RE) bounds govern stable recovery by norm-regularized estimators. For isotropic sub-Gaussian measurements, the benchmark sample size is $1+w(A)^2$, where $w(A)$ is the Gaussian width of the normalized descent cone. The COLT 2015 open-problem note (Banerjee et al., 2015) asked whether the same law...

Shi Fu, Hui-Bo Xu, Qi-Xin Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

On Basis Function Selection for Sparse Gaussian Process Regression

Sparse Gaussian processes achieve $O(N)$ inference by replacing the kernel with an appropriate expansion in a fixed basis $\{\phi_j\}$ on the input space. Given a compute budget $M \ll N$, practitioners conventionally truncate the basis to its first $M$ entries. Nothing in the formalism, however, prevents one from sele...

Marnix Van Soom, I. De Boi · 0 citations
Preprint Aug 2026

Stochastic gradient descent with initial regularization

A variant of stochastic gradient descent with initial regularization with initial regularization is analyzed and dimension-free upper bounds on its expected excess risk for the squared loss are derived.

Nabil Kahalé · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.