Skip to content
Preprint

The Sampling Distribution of the Log-Euclidean Distance Between Sample Correlation Matrices

Aug 2026 · 0 citations · 1 references
Mathematics

Abstract

Comparing correlation matrices across time or stress scenarios is critical in quantitative finance and multivariate statistics, yet sample estimation noise often obscures whether an observed distance reflects a true structural shift. We derive the asymptotic sampling distribution of the intrinsic off-log (log-Euclidean) distance between two independently estimated full-rank correlation matrices under the null hypothesis that their population correlation matrices coincide. Under general sampling with finite fourth moments, the scaled squared distance converges to a weighted sum of independent $\chi_1^2$ variables, with weights determined by the asymptotic covariance of the Generalized Fisher Transformation (GFT) coordinates. Under Gaussian sampling at independence, this simplifies to a parameter-free $4\chi_d^2$ law. To calibrate tail probabilities, we provide closed-form cumulant generating functions, Lugannani--Rice saddlepoint quantiles, and an explicit Chernoff envelope requiring no root-finding. The first moment of the limiting law establishes a simple rule of thumb for the baseline expected distance under the null hypothesis ($\operatorname E[d_{\mathrm{LE}}] \lesssim 2\sqrt{d/n}$ near independence), quantifying the average separation induced strictly by estimation error. We establish plug-in consistency, present an explicit Gaussian covariance factorization, compare the distance statistic with coordinate Wald tests, and characterize its local power.

View source

Similar papers

Preprint Aug 2026

Exact Likelihood and Sampling for Riemannian Gaussian Distributions on Correlation Matrices

Correlation matrices arise when marginal scales are removed from covariance matrices, yet a normalized likelihood must account for both quotient distance and quotient volume. We propose a Riemannian Gaussian model for full-rank correlation matrices under quotient-affine geometry. The distribution is proper and has finite radial moments. We derive exact score and profiled-scale equations and recover Fisher-transformed Gaussian inference for two-dimensional matrices. In higher dimension, a curvature calculation shows that the normalizing constant can vary with the center. Exact maximum likelihood and Fr\'{e}chet estimation may therefore have different population targets. We develop chart-based methods for evaluating the normalizer, fitting the likelihood, and sampling. Numerical studies verify the analytic case and compare integration, estimation, and sampling procedures across dimensions and dispersion regimes. A rolling-finance application and a controlled prior study illustrate both the value and computational cost of the model. The method is most reliable in small to moderate dimensions, while proposal efficiency and numerical conditioning deteriorate near the boundary and at larger dispersion.

Kisung You · 0 citations
Open access Aug 2026

Asymptotic Theory for Kernel Density Estimation Under Dependent Length-Biased Sampling

We establish an asymptotic theory for the Jones inverse-weighted kernel density estimator when length-biased observations form a strictly stationary short-range dependent sequence. The statistical difficulty is intrinsically composite: reciprocal weighting is singular at the origin, the normalizing mean is estimated from the same dependent sample, kernel localization shrinks with the bandwidth, and the centered summands form a row-wise stationary triangular array whose envelope diverges at rate hn−1. Under a non-negative compactly supported Lipschitz kernel, an inverse-moment condition, geometric α-mixing, local regularity of the target density, and uniform local bounds on lagged bivariate densities, we prove strong uniform consistency on compact subsets of (0,∞) and, separately, the uniform stochastic bound OP{hn2+(logn/(nhn))1/2}. A covariance-localization argument shows that the scaled serial-covariance contribution is O{hnlog(1/hn)}=o(1), so the first-order pointwise variance coincides with that of the corresponding independent length-biased estimator. Pointwise and finite-dimensional Gaussian limits are obtained by an explicit big-block/small-block argument with off-diagonal covariance control. The ratio normalization is treated directly: its variance contribution, its product with the localized fluctuation, and its cross-covariance with that fluctuation are all negligible at the nhn scale. We further derive first-order AMSE and AMISE criteria, their oracle bandwidths, and feasible pointwise studentization under undersmoothing. The numerical study separates oracle from data-driven bandwidth selection, evaluates full-ratio HAC and moving-block corrections, examines a Frank-copula Markov robustness design, and benchmarks the Jones estimator against an alternative length-biased estimator. The simulations support the first-order theory while demonstrating that persistent short-range dependence can remain consequential for finite-sample uncertainty.

Salim Bouzebda, S. Didi · 0 citations
Preprint Aug 2026

Connecting Riemannian Geometry and Statistical Inference for Correlation Matrices

The quotient-affine metric gives an intrinsic Riemannian geometry to full-rank correlation matrices, but its geodesic distance has no closed form and we are not aware of an analytic asymptotic null distribution for it. We connect this geometry, introduced in 2019, with Jennrich's 1970 asymptotic test for equality of correlation matrices. The quadratic form underlying Jennrich's statistic is exactly one half of the quotient-affine metric tensor. The identity arises because eliminating marginal standard deviations from Gaussian Fisher information performs the same projection as quotienting out diagonal rescalings. Jennrich's statistic therefore evaluates the local quotient-affine quadratic form directly. Moreover, for two independent Gaussian samples with a common population correlation matrix, the squared geodesic distance, scaled by effective sample size, converges in distribution to $4\chi^2_d$, where $d = p(p-1)/2$. For $p=2$, the result reduces to the two-sample Fisher $z$ test.

A. Kuketayev · 0 citations
Preprint Jul 2026

Restricted nonlinear shrinkage of high-dimensional residual covariance matrices in multivariate regressions

We study estimation of the p*p residual scatter (shape) matrix in a high-dimensional multivariate linear regression, where p and n grow proportionally. When the coefficient matrix obeys a known linear restriction of rank q<d, as in multivariate analysis of variance, growth-curve models, and reduced-rank regression, the restricted fit leaves additional residual degrees of freedom that sharpen estimation of the shape matrix. To accommodate heavy-tailed errors, we work with independent elliptically distributed rows under a mild scale condition, a finite second moment on the radii, which is far weaker than the usual sub-Gaussian assumptions and covers every multivariate-t law with more than two degrees of freedom. Shrinking the restricted residual sample covariance directly is unsound here, since its limiting spectrum depends on the radial distribution. We instead shrink a scale-invariant scatter of the restricted residuals, whose spectrum is distribution-free over the elliptical family and obeys the same limiting law as under Gaussian errors, at a smaller effective aspect ratio. The resulting estimator attains the rotation-equivariant oracle and is asymptotically optimal within that class, and a Stein-type combination with the unrestricted estimator dominates it while remaining safe under misspecification. We further correct for the case in which the restriction is itself selected from the data. Simulations, a growth-curve experiment, and two real-data analyses illustrate the results.

H. Karamikabir, Mohammad Arashi Department of Statistics, Faculty of Intelligent Systems Engineering et al. · 0 citations
Preprint Aug 2026

Approximating the null distribution of generalized distance covariance

The null distribution of distance covariance is usually approximated by permutation, which is prohibitive when very small p-values are needed, or by matching a few moments to a parametric family, which is inaccurate in the tails. A third option is to approximate the limiting distribution, a weighted sum of chi-square variables, directly through the spectra of the doubly centred distance matrices. This is used for kernel-based tests but has lacked a rigorous justification. We prove that the empirical spectra give a uniformly consistent approximation of the limiting null distribution, and hence an asymptotically valid test, for a general class of distances of negative type on separable metric spaces. The result covers the Hilbert-Schmidt independence criterion as a special case. We also give an adaptive algorithm that brackets the p-value from a partial eigendecomposition, reducing the cost from $O(n^3)$ to $O(k n^2)$, and a shrinkage correction matching the first two moments. In simulations, the proposed tests are the only non-Monte-Carlo procedures whose empirical type I error converges to the nominal level.

D. Edelmann · 0 citations
Preprint Jul 2026

Spectra of high-dimensional Spearman correlation matrices under scale-mixture dependence

We study the asymptotic spectral properties of high-dimensional Spearman correlation matrices for scale-mixture data. We consider observations of the form $x_t=\sigma_t \xi_t \in \mathbb{R}^N,$ where the coordinates of $\xi_t$ are i.i.d.\ and the scalar mixture variable $\sigma_t$ is shared by all coordinates. Under natural symmetry assumptions, the coordinates of $x_t$ are pairwise uncorrelated in both the Pearson and Spearman sense. Nevertheless, they are not independent when the mixture variable is non-degenerate. We show that this higher-order dependence survives the rank transformation and leaves a nontrivial spectral signature. In the proportional regime $N/T\to q\in(0,\infty),$ the empirical spectral distribution of the Spearman correlation matrix converges almost surely to a generalized Mar\v{c}enko--Pastur law governed by the limiting distribution of an effective rank variance. We also formulate a broader latent-variable extension, which covers, in particular, some scale-mixture models with correlated directional components. We discuss solvable examples and numerical approximations, motivated in part by heavy-tailed data in robust multivariate statistics, econometrics, and finance.

J. Bouchaud, Pierre Bousseyroux, Tomas Espana et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.