Classical canonical correlation analysis becomes numerically unstable when the number of variables is large relative to the sample size and is sensitive to contamination in observations or individual cells. This study develops an integrated robust and regularized procedure that combines bounded cellwise wrapping, shrinkage estimation of the joint correlation matrix, and robust reweighting in a low-dimensional canonical score space. The resulting observation weights enter a second regularized canonical correlation fit, so the final estimator remains well defined when the combined number of variables exceeds the sample size. The simulation study shows that relative estimation accuracy depends on the signal strength, contamination mechanism, and dimensional configuration. The proposed estimator is competitive in several moderate-signal settings and has a clear computational advantage, whereas the minimum regularized covariance determinant plug-in estimator provides lower estimation error in many high-signal configurations. An additional ultra-high-dimensional experiment demonstrates numerical feasibility with modest memory use but also reveals substantial attenuation, identifying a limitation of the present dense estimator. The results therefore support a regime-dependent interpretation rather than a claim of uniform superiority. The complete reproducible simulation workflow is provided.
In multivariate statistical analysis, accurate modeling of the covariance structure is critical for high-dimensional data analysis, variable selection, and regularization. In high-dimensional settings, strong inter-variable correlation and redundancy are key factors limiting the performance of classical sparsity-based methods. While LASSO and its variants provide effective tools for coefficient shrinkage and variable selection, they may select redundant variables and produce unnecessarily complex models in highly correlated settings. In this study, a Correlation-Sensitive Adaptive LASSO (CDA-LASSO) method is proposed to address these limitations. The proposed approach is based on a hybrid weighting mechanism that makes the penalty term sensitive not only to initial coefficient magnitudes but also to the correlation structure between variables. This structure incorporates correlation-based redundancy information and imposes stronger penalties on predictors with higher directed redundancy scores. Under fixed-dimensional regularity conditions, the bounded correlation multiplier is shown to preserve the selection consistency and oracle limiting distribution of Adaptive LASSO. The method was evaluated through 14 high-dimensional simulation scenarios covering different sample sizes, dimensionalities, sparsity levels, correlation strengths, support structures, and normal or heavy-tailed errors. The results indicate that the Max and kMean variants generally reduce the false discovery rate and model size relative to LASSO and Elastic Net while maintaining broadly comparable predictive performance. Numerical improvements over Adaptive LASSO were also observed in several scenarios, although these differences were not uniformly statistically significant. Under very high correlation, reductions in false discoveries were sometimes accompanied by modest decreases in the true positive rate. The real-world Riboflavin analysis further showed that the CDA-LASSO variants produced smaller models than LASSO and Elastic Net while retaining comparable prediction errors. Overall, CDA-LASSO directly incorporates the internal correlation structure of the data into the penalty weights without requiring a predefined graphical structure and provides a practical methodological extension for more controlled and parsimonious variable selection in high-dimensional correlated settings.
Y. Güral, Büşra Ceylan Kuzu, M. Gürcan· Symmetry· 0 citations
In the era of high-dimensional data, the classical assumption that the number of observations n vastly exceeds the number of variables p is frequently violated. When p and n grow proportionally (p/n → c > 0), the sample covariance matrix becomes severely distorted by sampling noise. Its eigenvalues are systematically biased: large population variances are overestimated, and small ones are underestimated. This phenomenon, governed by the Marchenko–Pastur law of Random Matrix Theory (RMT), renders standard statistical procedures highly unstable. This paper provides a comprehensive, mathematically rigorous treatment of spectral shrinkage, the optimal remedy for this distortion. We transition from the theoretical foundations of the Stieltjes transform to the practical implementation of rotationally invariant estimators. By combining formal proofs, geometrical interpretations, and reproducible R simulations with explicit console outputs, we demonstrate why spectral shrinkage is not merely a heuristic regularization technique, but a mathematically undeniable necessity for modern high-dimensional statistics.
Innocent Nsabimana· International Journal For Mu...· 0 citations
Classical MANOVA procedures are not directly applicable in high-dimensional settings where the number of variables is comparable to, or exceeds, the sample size, and many existing high-dimensional MANOVA tests remain sensitive to outlying observations. This study proposes a weighted minimum regularized covariance determinant (MRCD)-based robust Wilks’ Lambda test for one-way high-dimensional MANOVA. The proposed method combines MRCD-based robust location and scatter estimation with a robust distance-based reweighting step and uses permutation calibration to obtain p-values. Through extensive Monte Carlo simulations, the method is evaluated in terms of Type-I error control, power, and robustness under structured contamination. Under clean data, the proposed test maintains empirical Type-I error near the nominal level, with only modest aggregate differences from Cheng-GM; Schott’s test can have higher power under weak signals. Under contaminated null scenarios where outliers create artificial group separation, the proposed method yields lower false-rejection rates than the competitors considered. A controlled sensitivity illustration using breast-cancer gene-expression data shows the same qualitative behavior after imposed contamination. The method is therefore positioned as a robustness-oriented option for contamination-prone high-dimensional MANOVA, at the cost of additional computation.
This paper investigates the asymptotic behavior of the out-of-sample prediction risk of the high-dimensional ridgeless least-squares estimator when the feature dimension $p$ and the sample size $n$ grow proportionally. We consider a generalized spiked population covariance model with multiple latent factors, where the number of spiked eigenvalues may remain finite or increase with $n$, and the spiked eigenvalues may be bounded or diverge at arbitrary rates. Beyond characterizing the impact of covariance spectra, we reveal a new mechanism underlying benign overfitting: the prediction behavior of ridgeless interpolation is fundamentally governed by the alignment between the regression coefficient $\boldsymbol\beta$ and the spiked eigenspaces of the population covariance matrix. In particular, we show that the signal energy distributed along latent spike directions determines whether interpolation leads to benign, tempered, or catastrophic overfitting. Our theoretical framework establishes sharp prediction risk limits under minimal moment conditions, requiring only finite fourth moments rather than Gaussianity. We characterize how the number, strength, and geometric structure of the spikes jointly influence the double-descent phenomenon. These results provide a unified understanding of when latent covariance structures facilitate or hinder generalization in overparameterized regression.
We propose computationally efficient tests for equality of mean vectors of two or more high-dimensional populations. Central to our approach is an equivalence between equality of means and a zero population logistic regression parameter. We establish this equivalence for independently distributed observations without imposing common distributional assumptions across populations. Our procedure uses logistic Lasso to screen informative variables and an unpenalized logistic refit for inference in the reduced dimension, yielding asymptotically correct size and consistency. For a specified two-sample Gaussian submodel and sparse discriminative class, the test also attains the minimax separation rate. The framework extends to multiple populations through multi-class logistic regression. Simulations demonstrate accurate size control, strong power, and favorable computational scaling compared with existing tests under unbalanced designs and variance heterogeneity. Applications to gene-expression data with more than twenty-two thousand variables illustrate the practical scalability of the proposed procedures.
In this paper, we study the autocovariance matrix estimation and inference problems under heavy-tailedness, high-dimensionality, general nonlinear temporal dependence, and potentially nonstationarity of time series. We consider two types of tail-robust autocovariance matrix estimation methods: the element-wise Huber's $M$-estimator and a computationally more efficient element-wise truncated estimator. Both estimators are designed to achieve sharp error bounds in matrix max-norm. The nonasymptotic properties of these estimators are proved based on new variants of Bernstein-type inequalities under functional dependence for the potentially nonstationary processes which may be of independent interest. Moreover, we prove a high-dimensional Gaussian approximation result, as a limiting distribution, for our element-wise truncated autocovariance estimator. A Gaussian multiplier bootstrap result is also given to facilitate the practicality. Our theoretical results are nonasymptotic, which gives explicit error bounds in terms of the sample size, dimensionality, moments, and the strength of temporal dependence. Numerical evidence is provided to support our theoretical results. Finally, we illustrate the benefits of the proposed methodology for detecting change points in monthly macroeconomic data.
Hao-Tian Xu, S. Guerrier, Run-Ze Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.