Skip to content
Preprint

Sharp Minimax Theory for Randomized Experiments

Aug 2026 · 1 citation · 48 references
Mathematics Economics

Abstract

We study minimax-optimal designs and estimators for estimating the sample average treatment effect in finite population randomized experiments, where both design and estimator are unrestricted. For binary potential outcomes, we show this minimax risk is equivalent to the minimax risk $\rho_n^*$ of an estimation problem with $2$ unknown parameters. We leverage this reduction to establish a second-order risk expansion $\rho_n^* = n^{-1} - Cn^{-4/3} + o_n(n^{-4/3})$ for an explicit constant $C$ related to the Airy function. The minimax risk is attained by Bernoulli randomization with a nonlinear shrinkage estimator. Our results show that standard procedures such as complete randomization with difference in means are only minimax optimal up to first order in $n.$ We derive further results on admissibility of these procedures and discuss the practical implications of our results.

View source

Similar papers

Case report Open access Aug 2026

Optimal Experimental Design and Estimation when Potential Outcomes are Bounded

I study the optimal design and analysis of randomized experiments for estimating finite-population average treatment effects when potential outcomes are known to be bounded, as with binary outcomes. Among all assignment mechanisms and a broad class of affine estimators, worst-case mean-squared error (MSE) is minimized by independent random assignment and an unconventional regression of the support-midpoint-centered outcome on the recentered treatment, with no intercept. This contrasts with the usual prescription of balanced complete randomization and difference-in-means estimation: when outcomes are bounded, randomness in the realized treatment share is informative. The worst-case gain over full-sample complete randomization is asymptotically small, but gains can be first-order relative to other designs: complete within-pair randomization and pair-fixed-effect regression have twice the worst-case MSE. I extend the result to allow for arbitrary estimators. Independent random assignment remains optimal, and the generally-nonlinear optimal estimator can meaningfully reduce worst-case MSE.

Peter Hull · 0 citations
#data science Jul 2026

Simple-regret rates and minimax optimality of fixed-prior expected improvement in Matérn and squared-exponential RKHSs

We study the expected improvement (EI) policy for minimizing a deterministic objective function $f$ on a nonempty compact set $\mathcal X \subset\mathbb R^d$. We assume that $f$ belongs to the RKHS $\mathcal H_k$ of a continuous positive-semidefinite kernel $k$ on $\mathcal X$. Function values are observed exactly, and EI is computed from a fixed zero-mean Gaussian-process model with covariance $\sigma^2k$. After an initial design, the policy queries a point whose EI is at least a fixed positive fraction of its maximum. We identify the normalized posterior standard deviation at a candidate point $x$ with the norm of the corresponding innovation in the canonical feature space, namely the component of $k(x,\cdot)$ orthogonal to the span of the preceding evaluation representers. Sequential separation radii bound the ranked innovation norms along arbitrary query sequences. We estimate these radii using Gram determinants and Kolmogorov widths for subspaces of different dimensions, then combine the estimates with a one-step regret inequality to obtain finite-budget bounds for simple regret. After $N$ post-initial queries, simple regret is $O(N^{-\nu/d})$ for isotropic Mat\'ern kernels of smoothness $\nu>0$. For the isotropic squared-exponential kernel, simple regret is $O(\exp[-c_1\min\{N, N^{1/d}\log(eN)\}])$ for some $c_1>0$. With exact EI maximization, it is $O(\exp[-c_2N^{1/d} \log(eN)])$ for some $c_2>0$. For every fixed $B\geq0$, these bounds are uniform over the RKHS ball of radius $B$. If $\mathcal X$ has nonempty interior and $B>0$, then, among deterministic methods whose final recommendation may be any point of $\mathcal X$, the exact EI policy is minimax-rate optimal over the RKHS ball of radius $B$ for Mat\'ern kernels and minimax-rate optimal up to constants in the exponent for squared-exponential kernels.

Emmanuel Vazquez, S. Petit · 0 citations
Preprint Aug 2026

Gaussian-efficient testing by betting on the mean of bounded data

Given $[0,1]$-valued random variables $X_1,\dots,X_n$ such that $\mathbb{E}[X_i | X_1,\dots,X_{i-1}]= \mu$ for all $i$, we propose a new nonasymptotic confidence interval for $\mu$ that is obtained by inverting terminal e-values generated by a novel betting strategy. When the data are iid, its limiting width matches that of the central limit theorem (``Gaussian-efficient''), finally surpassing the inefficient limits of previous betting intervals. Our main conceptual advance involves designing betting fractions that track the conditional rejection probability of the most powerful terminal test in a limiting Gaussian experiment. When one predictable variance estimator is shared across candidate means, the deterministic inversion is an interval for every data sequence and its two endpoints can be found easily. The width can be improved further with external randomization. In simulations, our method yields the tightest intervals to date; for every distribution tested and all sufficiently large $n$, our deterministic version beats STaR-Bets and is competitive with Gaffke, while the randomized improvement beats both. It thus combines finite-sample validity under martingale dependence, easy endpoint computation, Gaussian-efficient inference for iid data, and excellent empirical performance. We also extend the construction and its efficiency theory to sampling without replacement, where it again achieves state-of-the-art empirical performance.

Diego Martinez-Taboada, Aaditya Ramdas · 0 citations
Preprint Aug 2026

The Limits of Experimental Design: Covariate Balance Beyond Low Dimension

We study how fast experimental designs can approach the semiparametric efficiency bound in finite samples, as measured by the excess variance of unadjusted treatment effect estimation. We prove an impossibility theorem: under weak conditions, no design can approach the variance bound uniformly over smooth outcome models unless covariate dimension $d \ll \log n$. Even in experiments with thousands of units, this permits only a handful of covariates. Motivated by this, we propose new designs based on discrepancy minimization that instead attempt to control imbalances over restricted-complexity nonparametric function classes. Such designs achieve fast rates to their corresponding restricted efficiency targets, permitting $d \ll n$ covariates in an additive nonparametric specification. They can also be combined with matching to protect against unmodeled outcome variation. In simulations calibrated to 12 published experiments, our designs reduce variance relative to matched pairs randomization in every empirical setting.

Max Cytrynbaum · 0 citations
Preprint Aug 2026

On the optimality of antithetic randomization for cross-validation

In the classical normal means problem, independent train--test folds can be constructed by perturbing the data with normal randomization. Averaging over $K$ such folds yields a cross-validation estimator whose bias depends on the marginal distribution of the randomization variables, while its variance depends on their joint distribution. This raises the questions: which joint law is optimal, and how to construct the corresponding randomization scheme? We show that: (i) for smooth estimators, antithetic randomization with pairwise correlation $\rho=-1/(K-1)$ is necessary and sufficient for the reducible variance due to randomization to remain bounded as the bias vanishes; (ii) a general construction yields a class of antithetic schemes, within which the jointly normal scheme is minimax optimal; and (iii) for non-smooth estimators with finitely many jump discontinuities, antithetic randomization improves the asymptotic rate of the reducible variance, while a simple control variate restores bounded variance when the discontinuities are known.

S. Chattopadhyay, Sifan Liu, Snigdha Panigrahi · 0 citations
Preprint Sep 2026

Finite-sample nonparametric mean tests: Leave-one-out duality and asymptotic optimality

We study finite-sample valid tests of the one-sided mean hypothesis $H_0:\mu\leq 1$ against $H_1:\mu>1$ for nonnegative random variables. To do so, we develop a leave-one-out dual certificate framework, where certain pointwise inequalities imply p-value validity under the conditional mean null $\mathbb{E}[X_i\mid\mathbf{X}_{-i}]\leq 1$, and which also gives conditions that allow combining dual certificates for p-values to show that their pointwise minimum is also a valid p-value. The framework proves finite-sample validity of Wang and Zhao's nonparametric likelihood-ratio statistic $T_{\mathrm{nplr}}$, yields a new p-value $T_{\mathrm{bin}+}$ extending the Clopper--Pearson binomial test to general nonnegative random variables, and shows that the pointwise minimum $\min\{T_{\mathrm{nplr}},T_{\mathrm{bin}+}\}$ is itself a valid and more powerful p-value. We establish sharp optimality results for such testing problems in two regimes: both $T_{\mathrm{nplr}}$ and $T_{\mathrm{bin}+}$ attain a universal detectability boundary for the null $H_0$ without moment or tail assumptions, and $T_{\mathrm{bin}+}$ attains a nonparametric power lower bound under $n^{-1/2}$-local alternatives to $H_0$. Efficient algorithms and numerical experiments demonstrate substantial finite-sample power gains over existing valid methods.

Yi-Fan Zhu, John C. Duchi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.