Skip to content
Preprint

Power and Limits of Subset Selection in Statistical Estimation

Jul 2026 · 0 citations · 18 references
Mathematics

Abstract

We study the power and limitations of subset selection in statistical estimation through the framework of \emph{super-teaching}, where a teacher selects a subset of i.i.d. data to optimize a learner's estimator. Unlike prior work focused on specific distributions or fixed subset sizes, we develop a general theory under minimal assumptions. For mean estimation, we prove that super-teaching is possible for any distribution whose density is bounded away from zero in some neighborhood of the mean, allowing subset sizes growing as $k = o(n^{1/3})$ and achieving error on the order of roughly $k!/n^{k}$. This significantly extends existing results on admissible distributions and subset scaling. We also extend the analysis to parameters expressed as smooth functionals of expectations, such as variance and scale parameters in classical parametric families, including settings with heavy tails. Moreover, we show that super-teaching can greatly improve estimation rates for nonlinear estimators like the sample median, achieving rates beyond classical asymptotics. Through examples, including cases where maximum likelihood estimators are inconsistent or fail to be asymptotically normal, we demonstrate that super-teaching can succeed even when standard statistical guarantees break down. Our results establish a unified theory of data selection to enhance statistical efficiency.

View source

Similar papers

Preprint Jul 2026

Local permutation tests for conditional independence: an adaptive binning perspective

In this work, we study the problem of testing conditional independence between random variables $X$ and $Y$ given a confounder $Z$. The local permutation test (LPT) offers a principled approach to this problem by partitioning the $Z$-space into pre-specified bins, and permuting the $X$ and $Y$ data within each bin, to assess the significance of an observed test statistic. However, when the partitions are pre-fixed, the resulting partition can be poorly balanced, as some bins may contain most of the samples while others contain only a few. This motivates the use of data-adaptive binning strategies, such as equisized bins with a fixed (typically small) number of points. We study this natural and practically important extension of LPT, providing finite-sample bounds on the Type I error for an arbitrary test statistic, providing stronger validity results than previously known. We also show that LPT attains power comparable to the oracle likelihood ratio tests derived from the Neyman-Pearson lemma. Within a linear confounder model class, we further analyze the effect of bin size and demonstrate that constant bin sizes can match the performance of partitions with growing bin-size. These results, further supported by extensive numerical simulations, position the proposed data-adaptive strategy as both practically implementable and statistically efficient.

David Chen, Rohan Hore, R. Barber · 0 citations
Preprint Jul 2026

Why Constants Matter in Distribution Testing: From Uniformity to Calibration

Distribution goodness-of-fit testing has developed a powerful rate-level theory: we often know how the required sample size scales with the alphabet size, the separation from the null, and the target error probability. Uniformity testing is the canonical example. One can distinguish the uniform distribution on $N$ categories from alternatives at total-variation distance at least $\epsilon$ with far fewer than $N$ samples, and the optimal scaling is now well understood. But rate-level theory leaves an important question unresolved: among several tests with the same sample-complexity order, which one actually gives the best risk or power? This is a constant-level question. It is especially relevant in modern applications where distribution testing is used not merely as an asymptotic abstraction, but as a practical design tool. This note argues that sharp constants in distribution testing play a role analogous to Fisher information in parametric estimation and Pinsker's constant in nonparametric estimation. First, they distinguish between tests that are all rate-optimal but not equally powerful. Second, they reveal the effective signal-to-noise ratio governing the testing problem. Third, they can guide tuning-parameter choices in downstream applications. We illustrate this perspective through large-alphabet uniformity testing and then explain why the same logic matters for choosing the number of bins in calibration testing.

A. Kipnis · 0 citations
Preprint Jul 2026

Testing the equality of estimable parameters

This paper proposes a general and unified framework for testing the equality of a broad class of parameters, defined via $U$-statistics, across multiple independent populations. This approach encompasses various common statistical problems, such as comparing variances, correlation coefficients, or Gini indices, among many others. We consider two test statistics, a Wald-type statistic and an ANOVA-type statistic. The asymptotic distribution of the first one is derived under a fixed-dimension regime, whereas the second one is studied under both fixed and increasing-dimension regimes, where the parameter dimension diverges with the sample size. Based on these limiting distributions, we construct test procedures enabling asymptotically exact inference without parametric assumptions. Additionally, an alternative null distribution estimator based on a weighted bootstrap approximation is studied, which is applicable to the ANOVA-type statistic under a fixed-dimension regime. The finite-sample performance and computational efficiency of the proposed procedures are evaluated through an extensive simulation study. Finally, an application to a real dataset illustrates the usefulness of the proposed methodology.

M. Romero-Madronal, M. R. Sillero-Denamiel, M. D. Jiménez–Gamero · 0 citations
Preprint Aug 2026

Distribution-free testing of linear type

We introduce a distribution-free goodness-of-fit test, termed the omega-1 test, which naturally complements the Kolmogorov--Smirnov test and Cram\'{e}r--von Mises test and can be viewed as their (piecewise) linear analog. Defined as an $\mathrm{L}^{1}$-functional of the empirical process, the test statistic improves on balancing sensitivity to localized and diffuse alternatives and gives a robust and interpretable measure of distributional discrepancy, apart from close connections to the Wasserstein 1-distance. For finite samples, we derive a finite-dimensional computational form for the statistic under general conditions, which leads to various explicit formulas for its null distribution. Under mild continuity assumptions, the limiting statistic is distribution-free, with explicit distribution formulas. In composite settings, the statistic is also compatible with the Khmaladze transformation, enabling asymptotically distribution-free testing. The limiting transformed statistic also has an explicit distribution that escapes reliance on intractable compensator processes or purely numerical evaluation. Simulation results indicate rapid convergence of the finite-sample distributions to their limiting counterparts and support the practical applicability of the test.

Wei-Xu Xia · 0 citations
Case report Open access Aug 2026

Regularized goodness-of-fit statistics and exact nonparametric confidence bands for distributions with application to household consumption

We study from a finite-sample viewpoint the problem of building tests and simultaneous confidence bands for cumulative distribution functions (CDFs), continuous or discrete. We emphasize procedures based on reweighted empirical distribution function (EDF) with shrinking bandwidths in the tails of the distribution. Since weighted statistics may have a problematic behavior when scaling factors are small (or zero), we propose to use regularized statistics. We consider a wide class set of modified EDF-type statistics, and give general characterizations of their distributions in the case of i.i.d. observations, so that the relevant distributions can be simulated in finite samples. We show that test criteria in the class studied can be implemented through the technique of Monte Carlo tests (MCT), so that the level is fully controlled in finite samples, irrespective of whether the tested distribution is continuous or discrete, without the need to establish an asymptotic distribution. Standard criteria such as the Kolmogorov-Smirnov (KS), Anderson-Darling (AD), Eicker (E), and Berk– Jones (BJ) statistics are covered as special cases. Confidence bands are built by inverting sup-type goodness-of-fit test statistics. We show that the bands based on regularized AD-type and E-type statistics have closed forms which are especially easy to compute, without nonlinear optimization. For continuous variables, the null distributions of the statistics do not depend on the CDF tested. For noncontinuous distributions, we show that the MCT approach transparently controls test levels irrespective of the distribution tested. We also establish monotonicity properties (based on nesting image sets) and a general dominance result, so continuous critical values are valid (conservative) critical points. We show in Monte Carlo simulations that the proposed regularized goodness-of-fit tests and confidence bands are numerically tractable, reliable and yield power and precision improvements over standard procedures. The proposed methods are applied to the distribution of households’ consumption in Kenya.

Jean-Marie Dufour, Mame Astou Diouf · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.