The triple shrinkage adaptive GO estimator is proposed, which extends the GO framework with adaptive, coefficient-specific weights, and produces flexible penalization and achieves oracle properties asymptotically, yielding performance comparable to methods that effectively know the true support.
In multivariate statistical analysis, accurate modeling of the covariance structure is critical for high-dimensional data analysis, variable selection, and regularization. In high-dimensional settings, strong inter-variable correlation and redundancy are key factors limiting the performance of classical sparsity-based methods. While LASSO and its variants provide effective tools for coefficient shrinkage and variable selection, they may select redundant variables and produce unnecessarily complex models in highly correlated settings. In this study, a Correlation-Sensitive Adaptive LASSO (CDA-LASSO) method is proposed to address these limitations. The proposed approach is based on a hybrid weighting mechanism that makes the penalty term sensitive not only to initial coefficient magnitudes but also to the correlation structure between variables. This structure incorporates correlation-based redundancy information and imposes stronger penalties on predictors with higher directed redundancy scores. Under fixed-dimensional regularity conditions, the bounded correlation multiplier is shown to preserve the selection consistency and oracle limiting distribution of Adaptive LASSO. The method was evaluated through 14 high-dimensional simulation scenarios covering different sample sizes, dimensionalities, sparsity levels, correlation strengths, support structures, and normal or heavy-tailed errors. The results indicate that the Max and kMean variants generally reduce the false discovery rate and model size relative to LASSO and Elastic Net while maintaining broadly comparable predictive performance. Numerical improvements over Adaptive LASSO were also observed in several scenarios, although these differences were not uniformly statistically significant. Under very high correlation, reductions in false discoveries were sometimes accompanied by modest decreases in the true positive rate. The real-world Riboflavin analysis further showed that the CDA-LASSO variants produced smaller models than LASSO and Elastic Net while retaining comparable prediction errors. Overall, CDA-LASSO directly incorporates the internal correlation structure of the data into the penalty weights without requiring a predefined graphical structure and provides a practical methodological extension for more controlled and parsimonious variable selection in high-dimensional correlated settings.
Y. Güral, Büşra Ceylan Kuzu, M. Gürcan· Symmetry· 0 citations
Classical canonical correlation analysis becomes numerically unstable when the number of variables is large relative to the sample size and is sensitive to contamination in observations or individual cells. This study develops an integrated robust and regularized procedure that combines bounded cellwise wrapping, shrinkage estimation of the joint correlation matrix, and robust reweighting in a low-dimensional canonical score space. The resulting observation weights enter a second regularized canonical correlation fit, so the final estimator remains well defined when the combined number of variables exceeds the sample size. The simulation study shows that relative estimation accuracy depends on the signal strength, contamination mechanism, and dimensional configuration. The proposed estimator is competitive in several moderate-signal settings and has a clear computational advantage, whereas the minimum regularized covariance determinant plug-in estimator provides lower estimation error in many high-signal configurations. An additional ultra-high-dimensional experiment demonstrates numerical feasibility with modest memory use but also reveals substantial attenuation, identifying a limitation of the present dense estimator. The results therefore support a regime-dependent interpretation rather than a claim of uniform superiority. The complete reproducible simulation workflow is provided.
Hasan Bulut, Müjgan Zobu, V. Saglam· Mathematics· 0 citations
We propose computationally efficient tests for equality of mean vectors of two or more high-dimensional populations. Central to our approach is an equivalence between equality of means and a zero population logistic regression parameter. We establish this equivalence for independently distributed observations without imposing common distributional assumptions across populations. Our procedure uses logistic Lasso to screen informative variables and an unpenalized logistic refit for inference in the reduced dimension, yielding asymptotically correct size and consistency. For a specified two-sample Gaussian submodel and sparse discriminative class, the test also attains the minimax separation rate. The framework extends to multiple populations through multi-class logistic regression. Simulations demonstrate accurate size control, strong power, and favorable computational scaling compared with existing tests under unbalanced designs and variance heterogeneity. Applications to gene-expression data with more than twenty-two thousand variables illustrate the practical scalability of the proposed procedures.
In real-world applications, data are often error-contaminated; naively applying conventional methods without accommodating the measurement error effects often yields inconsistent estimates. Biased results can be further exacerbated by the ultrahigh-dimensionality of covariates. Focusing on the widely used function-on-scalar linear regression model, this article develops new methods for simultaneous parameter estimation and variable selection with error-prone covariates that can be ultrahigh-dimensional. The proposed framework provides flexibility to handle different types of measurement error models. We rigorously establish asymptotic properties of the proposed estimators under mild conditions. Notably, the convergence rates and limiting distributions of the proposed estimators depend on the nature of measurement error. Our findings highlight the significant differences of settings with ultrahigh dimensions compared to scenarios with finite dimensions, as well as the drastically different influence of different measurement error processes. For efficient computation, we design algorithms with data-driven tuning. We evaluate the finite sample performance of the proposed method through simulation studies and a real data application, demonstrating its effectiveness in addressing the challenges posed by error-contaminated and ultrahigh-dimensional of covariates.
We study a class of Lasso based estimators obtained by applying a quadratic correction on the Lasso equicorrelation set. The penalty matrix determines both the magnitude and geometry of the correction and contains, among other cases, the isotropic Lasso--Ridge correction, least squares refitting, Gram proportional interpolation between the Lasso and least squares, and coordinate specific penalties. We first derive a closed form representation and isolate the positive gain component of the resulting prediction improvement. We then control the remaining stochastic linear term in expectation by localizing the random signed equicorrelation model around a deterministic reference support. This yields a finite sample expectation bound that explicitly accounts for the randomness induced by Lasso model selection. The resulting decomposition provides a unified framework for understanding when Lasso based quadratic corrections can improve prediction.
Guo Liu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.