Skip to content
Preprint

Empirical likelihood confidence regions for ordered bivariate means

Aug 2026 · 0 citations · 27 references
Mathematics

Abstract

Let $\boldsymbol{X}_i=(X_{1i},X_{2i})^\top$ be independent and identically distributed observations with mean $\boldsymbol{\mu}=(\mu_1,\mu_2)^\top$ constrained by $\mu_1\leq\mu_2$. We study empirical-likelihood inference for a fixed mean vector and distinguish it from the previously known test of equality against an ordered alternative. At a fixed interior point, the constrained empirical likelihood ratio has the usual $\chi^2_2$ limit. At a fixed boundary point $(m,m)^\top$, its limit is the chi-bar-square distribution $\tfrac12\chi^2_1+\tfrac12\chi^2_2$. By contrast, profiling the unknown common mean in the equality-versus-order test yields $\tfrac12\chi^2_0+\tfrac12\chi^2_1$, the $k=2$ ordered-mean case of El Barmi (1996). We give an exact reduction of the latter statistic to the empirical likelihood of the paired differences, establish the localization step needed for the fixed-boundary expansion, and derive a local-to-boundary limit showing that interior calibration is not uniform over $n^{-1/2}$-neighborhoods of the boundary. Monte Carlo experiments under Gaussian, Student $t_5$, and shifted log-normal sampling examine fixed, boundary, and local regimes with explicit numerical-failure accounting. Illustrative paired-data analyses show the practical distinction between fixed-candidate confidence regions, directional equality tests, and ordinary scalar empirical-likelihood intervals truncated to the nonnegative parameter space.

View source

Similar papers

Jul 2026

Gaffke's confidence interval for the mean of bounded data is inadmissible but asymptotically efficient

Given observations $\mathbf x=(x_1,\dots,x_n)$, Gaffke (2005) defined \[ K_n(\mathbf x)=\mathbb{P}_{\mathbf D}\!\left\{\sum_{i=1}^n x_iD_i\le 1\right\}, \qquad (D_0,D_1,\ldots,D_n)\sim\mathrm{Dirichlet}(1,\ldots,1), \] and conjectured that it is a $p$-value whenever the inputs are independent e-values. Recently, Vlassis and Thomas (2026) proved this conjecture. Inverting the tests for observations in $[0,1]$ gives the confidence interval studied by Learned-Miller and Thomas (2020), which reduces to Clopper--Pearson for Bernoulli data. We give a finite- and large-sample account of Gaffke's test and interval. First, for every $\mathbf x\in[0,\infty)^n$ and every elementary symmetric polynomial $e_k$, \( K_n(\mathbf x)e_k(\mathbf x)\le {n\choose k}, \) so the Gaffke $p$-value never larger than the SymPol $p$-value of Ming et al. (2026). However, Gaffke's p-value is inadmissible. For $n=2$, we construct a valid rule that is strictly smaller on mixed configurations and is the unique admissible rule that dominates $K_2$. A neutral-face extension proves inadmissibility of $K_n$ for every $n\ge2$. If one independent uniform random variable is allowed, there is an even simpler full-dimensional improvement: on the upper orthant, where $K_n(\mathbf x)=1/\prod_i x_i$, replace it by $U/\prod_i x_i$. The equal-tail Gaffke confidence interval $I_n$ is nevertheless first-order asymptotically efficient: for iid observations on $[0,1]$ with unknown variance $\sigma^2>0$, \[ \sqrt n\,\operatorname{Width}(I_n)\longrightarrow 2\sigma z_{1-\alpha/2}\qquad\text{almost surely}. \] Our simulations also find that, among a variety of bounded-mean intervals considered, the Gaffke interval is the shortest, including comparisons with a recent empirical Berry--Esseen procedure having the same first-order Gaussian target.

Jiahao Ming, Aaditya Ramdas, Yi Shen et al. · 4 citations · ⚡1
Preprint Aug 2026

Tensor-normal maximum likelihood estimation at the operator-norm sample threshold

Let $X_1,\ldots,X_n$ be independent Gaussian tensors in $\mathbb{R}^{d_1}\otimes\cdots\otimes\mathbb{R}^{d_k}$ with a common covariance matrix given by the Kronecker product of $k$ unknown positive-definite factors, and let $D=\prod_{a=1}^k d_a$ and $d_{\max}=\max_a d_a$. Franks et al. (2026) established condition-number-free guarantees for the tensor-normal maximum likelihood estimator under the sample-size condition $nD\gtrsim k^2 d_{\max}^3$ and asked whether the cubic dependence on $d_{\max}$ could be reduced to a quadratic one. We answer this question affirmatively. For $t\geq 1$, if $nD\geq C k^2 d_{\max}^2 t^2$, then with high probability the maximum likelihood estimator exists, is unique, and satisfies $d_{\rm FR}(\widehat\Theta,\Theta)\leq C t \sqrt{k} d_{\max}/\sqrt{n}$ and $d_{\rm FR}(\widehat\Theta_a,\Theta_a)\leq C t\sqrt{k d_a} d_{\max}/\sqrt{nD}$ for every mode $a$. For every mode $a$ with $d_a=d_{\max}$, we further establish the sharp Thompson-metric bound $d_{\rm op}(\widehat\Theta_a,\Theta_a)\leq C t d_{\max}/\sqrt{nD}$. These guarantees are uniform over the unknown covariance factors and require neither condition-number bounds nor sparsity assumptions. Gaussian submodel lower bounds match the full and largest-factor Fisher--Rao rates up to a factor of $\sqrt{k}$ and the largest-factor Thompson rate up to universal constants. Consequently, for fixed $k$, the quadratic dependence of the sample-size threshold on $d_{\max}$ is optimal. GPT-5.6 Sol and Claude Fable 5 were used to assist with proof development, verification, and manuscript preparation.

Hengzhi He, Guang Cheng · 0 citations
Jul 2026

Estimating eigenvectors and eigenspaces of covariance matrices: Optimal Bounds and Conditions for Consistency

Let $X = [ \xi_1, \,\, \xi_2,...\,\, ,\xi_d]^\top$ be a zero-mean random vector of large dimension $d$ ($d \rightarrow \infty$) with (hidden) covariance matrix $M = (m_{ij})_{1 \leq i, j \leq d},$ where $m_{ij} = m_{ji} = \textbf{Cov}(\xi_i, \xi_j).$ Let $X_1, X_2, \dots, X_n$ be $n$ iid samples of $X$. Consider the sample covariance matrix $$\textstyle \tilde{M} := \frac{1}{n} \sum_{i=1}^{n} X_i X_i^\top.$$ In practice, one frequently uses the eigenvectors and eigenspaces of $\tilde M$ as estimators for those of $M$. A central task is to provide an error analysis for these estimators. In this paper, we provide an optimal error analysis, obtaining upper and lower bounds of matching order of magnitude, for a wide range of parameters $d$ and $n$, under mild assumptions on $M$. As corollaries, we obtain new necessary and sufficient conditions for the consistency of the estimators. In these conditions, we only require the number of samples $n$ to depend linearly on the effective rank of $M$, which can be much smaller than the dimension $d$.

Phuc Tran, Van H. Vu · 0 citations
Preprint Aug 2026

Sharp Tail Bounds Beyond Twice the Mean

Consider $n$ independent, non-negative, mean at most one random variables, $X_1,X_2,\ldots$. We show the following bound on the probability of their sum exceeding a threshold $t$: \[ \mathbb{P}\left[\sum_{i=1}^n X_i\ge t\right] \leq 1-\left(1-\frac{1}{t}\right)^n \text{ for all } t\ge 2n+1 \,. \] To prove this, we consider a relaxed optimization problem over a set of sequences of ordered, but non-independent random variables. This allows us to reformulate it recursively as dynamic programming problem. The bound becomes an equality for the binary i.i.d.~random variables satisfying $\mathbb{P}\left[X_i=0\right]= 1-\frac{1}{t}$ and $\mathbb{P}\left[X_i=t\right]=\frac{1}{t}$, which remains the maximizer in the relaxed problem.

P. Strack, Jannik M. Westermann · 1 citation
Preprint Jul 2026

A Simple Example of Bayesian Nonparametric Inconsistency

I present a very simple example in which a full-support prior over distribution functions has an inconsistent posterior, which I believe has instructive value. The example is $Y_i \stackrel{\text{iid}}{\sim} P_0$ under the Bayesian hierarchical model $[Y_i \mid P] \stackrel{\text{iid}}{\sim} P$ and $P \sim \int \mbox{DP}(\alpha, H) \, \pi_\alpha(\alpha) \ d\alpha$ when $\pi_\alpha(\cdot)$ has exponential (or heavier) tails, where $\mbox{DP}(\alpha, H)$ denotes a Dirichlet process with concentration parameter $\alpha$ and mean $H(\cdot)$; on the other hand, consistency is obtained under a light-tailed prior. This unifies and generalizes a consistency result of Freedman and Diaconis (1983) with an inconsistency result described by Ferguson et al. (1992).

A. Linero · 0 citations
Preprint Aug 2026

Absolute continuity of two-dimensional polynomial random vectors

Let $X=\{X_j\}_{j=1}^\infty$ be a sequence of independent random variables whose densities and moments of order $2d$ are uniformly bounded. For a random vector $f(X)=(f_1(X),f_2(X))$ whose components are polynomial functionals of degree at most $d$, we prove that \[ [[f]]_{\mu,\infty}^{\frac1{2d-1}}\mu(f\in A) \le C\bigl(\lambda_2(A)\bigr)^{\frac1{2d-1}} \] for every Borel set $A\subset\mathbb R^2$, where $C$ depends only on $d$ and the uniform density and moment bounds, and $\lambda_2$ denotes the Lebesgue measure on $\mathbb R^2$. Here $[[f]]_{\mu,\infty}$ measures the failure of proportionality of the highest-order orthogonal-chaos components of $f_1$ and $f_2$ with respect to the law $\mu$ of $X$. Consequently, whenever these components are not proportional, the law of $f$ admits a density in the weak Lorentz space $L^{\frac{2d-1}{2d-2},\infty}(\mathbb R^2)$. This recovers the dichotomy established by Nualart and Tudor for two-dimensional Wiener chaos vectors and extends it beyond the Gaussian setting. We also obtain the lower bound \[ \int_{\mathbb R^\infty}\Delta_f\,d\mu \ge C[[f]]_{\mu,\infty}^2, \] where $\Delta_f$ is the determinant of the Gram matrix of $\nabla f_1$ and $\nabla f_2$. In the special case of Gaussian measures, this gives a relaxed version of the estimate conjectured by Nourdin, Nualart, and Poly.

Egor D. Kosov · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.