Skip to content
Preprint

Empirical Bayes linear regression in high dimensions: Method of moments and sub-linear sample complexity

Aug 2026 · 0 citations
Mathematics

TL;DR

Under mild design conditions, satisfied by a broad class of correlated random designs, it is shown that EBMoM consistently estimates a growing number of moments and hence the prior itself, provided that $n\geq p^{1-o(1)}$.

Abstract

We study empirical Bayes estimation of the prior in high-dimensional linear regression $\mathbf{y}=\mathbf{X}\mathbf{\beta}+\mathbf{\varepsilon}$, where the regression coefficients are drawn independently from an unknown sub-Gaussian prior. In contrast to the sequence model, the design matrix couples the latent coefficients, so that recovering the prior requires deconvolving it from both the noise and copies of itself. We introduce the \emph{Empirical Bayes Method of Moments} (EBMoM), a computationally efficient procedure for general designs that recursively estimates the prior moments through a lower-triangular system of estimating equations and runs in time $O(np^2)$. Under mild design conditions, satisfied in particular by a broad class of correlated random designs, we show that EBMoM consistently estimates a growing number of moments and hence the prior itself, provided that $n\geq p^{1-o(1)}$. A matching information-theoretic lower bound, valid for a broad class of designs, shows that this sub-linear sample complexity is optimal for nonparametric prior estimation. This improves on existing results for likelihood-based methods whose consistency requires a linear sample size $n=\Omega(p)$.

View source

Similar papers

Preprint Sep 2026

Improved Variance Estimation in Homoskedastic Nonparametric Random-Design Regression via a Two-Scale Approach

We study estimation of a constant conditional variance $\sigma^2$ in nonparametric regression with a $d$-dimensional random design. This is an important problem, and similar questions arise in causal inference. The regression function is $\beta_b$-H\"older smooth, the design density is $\beta_g$-H\"older smooth and bounded above and away from zero, and we consider the nonparametric regime $\beta_b>1$ and $d>4\beta_b$. Set $\beta_g^\star=\beta_b(1-4\beta_b/d)/\{1+2\beta_b/d+8(\beta_b/d)^2\}$. We give an estimator whose mean squared error is upper bounded by $Cn^{-4(\beta_b+1)/(d+4)}$ in the low-regularity regime when $0<\beta_g\leq\beta_g^\star$. The low-regularity branch is based on a new two-scale construction: the covariate space is partitioned into cells, the local polynomial trend is projected out within each suitable cell, and the squared normalized contrast from one eligible close pair per cell is averaged across cells. In the high regularity regime when $\beta_g>\beta_g^\star$, a higher-order influence function estimator of Robins, Li, Tchetgen Tchetgen, and van der Vaart (2008) provides the rate $Cn^{-8\beta_b/(d+4\beta_b)}$. We also give an all-pairs ridge extension, which achieves the same two-scale rate, and evaluate the methods alongside a range of existing estimators in simulations.

Edgar Dobriban, Rajarshi Mukherjee, James M. Robins et al. · 0 citations
Preprint Jul 2026

Exact Generalization Error Curves of Kernel Ridge Regression for Functional Moment Estimation

Kernel ridge regression is a standard method for functional data analysis, but its exact behavior is less understood. We study tensor-product kernel ridge regression for estimating the $r$-th moment function of a random function based on noisy discrete observations. The formulation includes mean estimation, covariance estimation, and higher-order moment estimation in a single framework. Our main result gives a precise $1+o_{\mathbb{P}}(1)$ expansion for the $L^2$ error at each admissible regularization parameter. The expansion consists of bias and three variance terms corresponding respectively to variation across the independent sample paths, latent signal variation at each sample point, and variation from measurement errors, identifying the refined error structure underlying functional data. As applications, we show that KRR attains the minimax rate for source smoothness $s \leq 2$ but becomes suboptimal in the sparse regime for $s>2$ due to saturation. A technical ingredient is a set of concentration inequalities for $U$-statistics suited to the dependent product structure of functional observations.

Yi Ding, Yicheng Li · 0 citations
Preprint Sep 2026

Finite-sample nonparametric mean tests: Leave-one-out duality and asymptotic optimality

We study finite-sample valid tests of the one-sided mean hypothesis $H_0:\mu\leq 1$ against $H_1:\mu>1$ for nonnegative random variables. To do so, we develop a leave-one-out dual certificate framework, where certain pointwise inequalities imply p-value validity under the conditional mean null $\mathbb{E}[X_i\mid\mathbf{X}_{-i}]\leq 1$, and which also gives conditions that allow combining dual certificates for p-values to show that their pointwise minimum is also a valid p-value. The framework proves finite-sample validity of Wang and Zhao's nonparametric likelihood-ratio statistic $T_{\mathrm{nplr}}$, yields a new p-value $T_{\mathrm{bin}+}$ extending the Clopper--Pearson binomial test to general nonnegative random variables, and shows that the pointwise minimum $\min\{T_{\mathrm{nplr}},T_{\mathrm{bin}+}\}$ is itself a valid and more powerful p-value. We establish sharp optimality results for such testing problems in two regimes: both $T_{\mathrm{nplr}}$ and $T_{\mathrm{bin}+}$ attain a universal detectability boundary for the null $H_0$ without moment or tail assumptions, and $T_{\mathrm{bin}+}$ attains a nonparametric power lower bound under $n^{-1/2}$-local alternatives to $H_0$. Efficient algorithms and numerical experiments demonstrate substantial finite-sample power gains over existing valid methods.

Yi-Fan Zhu, John C. Duchi · 0 citations
Preprint Aug 2026

Beyond Modern Asymptotics for Log-Likelihood Ratios in Logistic Regression

We characterize the finite sample behavior of the log-likelihood ratio statistic in binary logistic regression, uniformly over both the design and the target parameter. For $n\geq d\geq 3$, we determine, up to universal constants, its worst case $(1-\delta)$ quantile over all fixed collections of design vectors and all target parameters: \[ d\log\left(\frac{e n}{d}\right)+\log\left(\frac{1}{\delta}\right). \] This is a nonasymptotic analogue of the Wilks $\chi^2_d$ phenomenon and requires no regularity assumptions on the design. The low dimensional cases exhibit unusual behavior. The worst case quantile in dimension $d=2$ is sharply of order \[ \log\log\log n+\log\left(\frac{1}{\delta}\right). \] The worst case quantile in dimension $d=1$ is of order $\log(1/\delta)$, with no dependence on $n$. Finally, i.i.d. Gaussian design vectors recover the classical Wilks scale. In the regime $n\gtrsim d+\log(1/\delta)$, we prove the sharp bound \[ d+\log\left(\frac{1}{\delta}\right). \] Unlike existing asymptotic results, our bounds are uniform over the target parameter, which may depend on $n$, $d$, and $\delta$.

H. Chardon, R. Pathak, Nikita Zhivotovskiy · 0 citations
Preprint Aug 2026

Non-Gaussian fluctuations for traces of squared sample correlation matrices in high dimensions

We provide limit theory for the trace of the squared sample correlation matrix $\mathbf R$, constructed from $n$ observations of a $p$-dimensional random vector with iid components. If the entries have finite fourth moment and $p$ and $n$ grow proportionally, it is known that $\operatorname{tr}({\mathbf R}^2)$ satisfies a central limit theorem (CLT) and the centering and scaling sequences are universal in the sense that they do not depend on the entry distribution. Under symmetry and regular variation assumption with index $\alpha$ and any growth rate of the dimension, we prove that the universal CLT remains valid for $\alpha>3$. For $\alpha<3$, we identify a critical dimension growth at which the fluctuations of $\operatorname{tr}({\mathbf R}^2)$ become non-Gaussian. Moreover, if the dimension $p$ grows faster and $\alpha\le 3$ we establish a non-universal CLT with norming sequences depending on the value of $\alpha$. Our findings are illustrated in a simulation study.

J. Heiny, Xuechun Hu, Felix J. Seo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.