Skip to content
Preprint

Optimal use of a black-box learner in semiparametric estimation

Jul 2026 · 1 citation · ⚡ 1 influential · 56 references
Mathematics

Abstract

Consider the partial linear model $Y = \mu_0(X) + \beta_0 \cdot T + \varepsilon$ and $T = \pi_0(X) + u$ in the structure-agnostic setting, where we are blind to the structure $\mu_0$ and $\pi_0$ and estimate the nuisances by a black-box hypothesis class. The learnability of the class is characterized by the estimation error $\delta_s$ in the absence of model misspecification and its $L_2$ mis-specification error $\delta_{a, \mu}$ and $\delta_{a, \pi}$ for $\mu_0$ and $\pi_0$, respectively. We propose a novel estimator of the target linear coefficient $\theta_0 = \beta_0$ with error rate \[ \frac{1}{\sqrt{n}} + \delta_{a, \mu} \cdot \delta_{a, \pi} + [\delta_s]^2. \] A matching lower bound is also established, implying that this rate is unimprovable. Compared with the product rate yielded by double machine learning (DML), our estimator removes the suboptimal term $\max(\delta_{a, \mu}, \delta_{a, \pi})\cdot \delta_s$ at no extra cost or assumption. Building on the underlying insights, which are neither tailored to the one-learner setting nor the partial linear model, we propose Transductive Adversarial Moment-calibrated Editing (TAME), which locally edits debiasing weights induced by black-box regression estimates on the inference sample through adversarial conditional moment calibration. TAME can be combined with any initial black-box estimates and can strictly improve on DML guarantees when the nuisance difficulties are imbalanced. We discuss how to fully exploit the advantages introduced by TAME, including the gains from using two learners, the resulting under-smoothing principle for model selection, and extensions to other linear functional estimation problems.

View source

Similar papers

Preprint Aug 2026

An Optimal Agnostic PAC Algorithm

This paper settles the sample complexity of agnostic PAC learning up to universal constants at every fixed $L^*$, matching the lower bounds of Devroye, Gyorfi, and Lugosi.

Markus Engelund Mathiasen, Jian Qian, Nikita Zhivotovskiy · 0 citations
Preprint Sep 2026

Improved Variance Estimation in Homoskedastic Nonparametric Random-Design Regression via a Two-Scale Approach

We study estimation of a constant conditional variance $\sigma^2$ in nonparametric regression with a $d$-dimensional random design. This is an important problem, and similar questions arise in causal inference. The regression function is $\beta_b$-H\"older smooth, the design density is $\beta_g$-H\"older smooth and bounded above and away from zero, and we consider the nonparametric regime $\beta_b>1$ and $d>4\beta_b$. Set $\beta_g^\star=\beta_b(1-4\beta_b/d)/\{1+2\beta_b/d+8(\beta_b/d)^2\}$. We give an estimator whose mean squared error is upper bounded by $Cn^{-4(\beta_b+1)/(d+4)}$ in the low-regularity regime when $0<\beta_g\leq\beta_g^\star$. The low-regularity branch is based on a new two-scale construction: the covariate space is partitioned into cells, the local polynomial trend is projected out within each suitable cell, and the squared normalized contrast from one eligible close pair per cell is averaged across cells. In the high regularity regime when $\beta_g>\beta_g^\star$, a higher-order influence function estimator of Robins, Li, Tchetgen Tchetgen, and van der Vaart (2008) provides the rate $Cn^{-8\beta_b/(d+4\beta_b)}$. We also give an all-pairs ridge extension, which achieves the same two-scale rate, and evaluate the methods alongside a range of existing estimators in simulations.

Edgar Dobriban, Rajarshi Mukherjee, James M. Robins et al. · 0 citations
#machine learning Preprint Sep 2026

Sharp Structure-Agnostic Minimax Risk for Partial Linear Models

We characterize the sharp structure-agnostic minimax risk for coefficient estimation in the partial linear model when the outcome and treatment nuisances are learned by two distinct black-box learners, which resolves the open problem in double machine learning posed by Gu (2025). For each nuisance \(q\in\{\mu,\pi\}\), we characterize the available learner by an approximation-error budget \(a_q\) and a stochastic-error budget \(s_q\), with the latter controlled through localized Rademacher complexity. Writing \(\mathcal E_n\) for the minimax mean-squared error, we show that \[\mathcal E_n\asymp1\wedge\left\{\frac1n+\left(a_\mu a_\pi+\min\left\{a_\pi s_\mu+s_\pi^2,\,a_\mu s_\pi+s_\mu^2\right\}\right)^2\right\}.\] The main new ingredient is a novel lower bound for the general two-learner problem. Our proof constructs four finite-mixture testing experiments using orthogonal code functions. Across these experiments, the hidden perturbations are placed outside both learner classes, outside only the treatment learner class, outside only the outcome learner class, or inside both learner classes. These four configurations capture, respectively, the interaction between the two approximation errors, the two asymmetric interactions between one learner's approximation error and the other learner's learning error, and the joint estimation difficulty of learning both nuisances. Combining the four resulting lower bounds yields the displayed rate, which matches the latest upper bound in Gu (2026). Our result shows that standard double machine learning can overstate the intrinsic difficulty of target estimation and provides a target-specific principle for learner selection: approximation error and stochastic complexity must be jointly balanced across the two nuisance learners rather than optimized separately.

Hai-Chen Hu, David Simchi-Levi · 0 citations
Preprint Aug 2026

Empirical likelihood confidence regions for ordered bivariate means

Let $\boldsymbol{X}_i=(X_{1i},X_{2i})^\top$ be independent and identically distributed observations with mean $\boldsymbol{\mu}=(\mu_1,\mu_2)^\top$ constrained by $\mu_1\leq\mu_2$. We study empirical-likelihood inference for a fixed mean vector and distinguish it from the previously known test of equality against an ordered alternative. At a fixed interior point, the constrained empirical likelihood ratio has the usual $\chi^2_2$ limit. At a fixed boundary point $(m,m)^\top$, its limit is the chi-bar-square distribution $\tfrac12\chi^2_1+\tfrac12\chi^2_2$. By contrast, profiling the unknown common mean in the equality-versus-order test yields $\tfrac12\chi^2_0+\tfrac12\chi^2_1$, the $k=2$ ordered-mean case of El Barmi (1996). We give an exact reduction of the latter statistic to the empirical likelihood of the paired differences, establish the localization step needed for the fixed-boundary expansion, and derive a local-to-boundary limit showing that interior calibration is not uniform over $n^{-1/2}$-neighborhoods of the boundary. Monte Carlo experiments under Gaussian, Student $t_5$, and shifted log-normal sampling examine fixed, boundary, and local regimes with explicit numerical-failure accounting. Illustrative paired-data analyses show the practical distinction between fixed-candidate confidence regions, directional equality tests, and ordinary scalar empirical-likelihood intervals truncated to the nonnegative parameter space.

N. Garg · 0 citations
Preprint Aug 2026

The Optimal Discounting Parameter of the Power Prior under Predictive Log-Loss

The power prior of Ibrahim and Chen incorporates historical data into a Bayesian analysis by raising the historical likelihood to a power $a_0 \in [0, 1]$. The choice of the exponent has remained an open question. This paper gives a closed-form answer under the predictive log-loss. For a model with $d$ parameters, a historical sample of size $N_0$, and average Kullback--Leibler divergence $\bar{D}_0$ between the historical and current data-generating distributions, the optimal exponent is $a_0^{*} = d/(2 N_0 \bar{D}_0 + d)$. Equivalently, the optimally borrowed effective sample size obeys the harmonic law $1/E^{*} = 1/N_0 + 2\bar{D}_0/d$: compatible data are pooled in full, and any difference caps the borrowed information at $d/(2\bar{D}_0)$ observations. The result is exact for multinomial data and extends to smooth parametric families. The law benchmarks adaptive borrowing, explains the reported degeneracy of the normalized power prior, and shows that neither subsetting the data nor decaying the exponent improves on the correctly discounted constant.

Yuriy A. Reznik · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.