Aug 2026· Entropy· Vol 28, pp. 933· 0 citations· 53 references
Medicine
TL;DR
A framework that evaluates the Sharma–Mittal entropy volumetrically across a two-dimensional parameter region rather than for a single parameter pair is proposed, demonstrating that volumetric entropy metrics defined across the entire parameter space provide a feature-ranking tool that is independent of parameter selection for continuous variables.
Abstract
Feature selection is a critical step in regression problems where a large number of continuous explanatory variables explain the same target through different dependency structures. Classical filters may remain sensitive to a single form of dependence, a single scale, or a specific discretization scheme; generalized entropy measures, on the other hand, typically require the parameters to be fixed at a single point. This study proposes a framework that evaluates the Sharma–Mittal entropy volumetrically across a two-dimensional parameter region rather than for a single parameter pair. For the continuous target and explanatory variables, the marginal, joint, and conditional densities are obtained using a Gaussian kernel density estimation; the conditional entropy and information gain surfaces are integrated across the region Ω = [0.05, 0.95]2 in the α-β plane to define three indices: PICSME, which measures the conditional uncertainty volume; PIGSME, which measures the gain volume; and NIGSME, which is the ratio of this gain to the total entropy volume of the target. The method is supported by bandwidth consistency and the renormalization of conditional densities; thus, the issue of negative gain that can occur in the continuous variables is resolved, yielding positive and interpretable scores across all six datasets. It is formally demonstrated that the fact that the three indices produce the same ranking is not an empirical observation but rather the result of a monotonicity relationship valid under a fixed target entropy volume. The method is compared with Pearson and Spearman correlations, the Shannon information gain, mutual information, and random forest variable importance across six regression datasets (Airfoil Self-Noise, AirQualityUCI, BodyFat, Meteorology, Concrete, and WineQualityWhite) that differ in their sample size, dimensions, and application domain. The evaluation is not limited to ranking consistency; the out-of-sample prediction performance is measured using least-squares models on the top-k subsets, with rankings calculated from the training partition. The findings show that NIGSME exhibits a performance comparable to that of built-in filters, outperforms them on the Concrete and Meteorology datasets, and never ranks as the weakest method on any dataset. The results demonstrate that volumetric entropy metrics defined across the entire parameter space provide a feature-ranking tool that is independent of parameter selection for continuous variables.
In multivariate statistical analysis, accurate modeling of the covariance structure is critical for high-dimensional data analysis, variable selection, and regularization. In high-dimensional settings, strong inter-variable correlation and redundancy are key factors limiting the performance of classical sparsity-based methods. While LASSO and its variants provide effective tools for coefficient shrinkage and variable selection, they may select redundant variables and produce unnecessarily complex models in highly correlated settings. In this study, a Correlation-Sensitive Adaptive LASSO (CDA-LASSO) method is proposed to address these limitations. The proposed approach is based on a hybrid weighting mechanism that makes the penalty term sensitive not only to initial coefficient magnitudes but also to the correlation structure between variables. This structure incorporates correlation-based redundancy information and imposes stronger penalties on predictors with higher directed redundancy scores. Under fixed-dimensional regularity conditions, the bounded correlation multiplier is shown to preserve the selection consistency and oracle limiting distribution of Adaptive LASSO. The method was evaluated through 14 high-dimensional simulation scenarios covering different sample sizes, dimensionalities, sparsity levels, correlation strengths, support structures, and normal or heavy-tailed errors. The results indicate that the Max and kMean variants generally reduce the false discovery rate and model size relative to LASSO and Elastic Net while maintaining broadly comparable predictive performance. Numerical improvements over Adaptive LASSO were also observed in several scenarios, although these differences were not uniformly statistically significant. Under very high correlation, reductions in false discoveries were sometimes accompanied by modest decreases in the true positive rate. The real-world Riboflavin analysis further showed that the CDA-LASSO variants produced smaller models than LASSO and Elastic Net while retaining comparable prediction errors. Overall, CDA-LASSO directly incorporates the internal correlation structure of the data into the penalty weights without requiring a predefined graphical structure and provides a practical methodological extension for more controlled and parsimonious variable selection in high-dimensional correlated settings.
Y. Güral, Büşra Ceylan Kuzu, M. Gürcan· Symmetry· 0 citations
We consider dimensionality reduction for high-dimensional observations accompanied by a supplied partition into two or more clusters. The objective is not to construct a low-rank projection, but to retain an interpretable subset of the original coordinates that preserves the distributional information distinguishing the clusters. For each coordinate, the proposed procedure compares the cluster-specific empirical distribution functions through a several-sample Kolmogorov-Smirnov separation statistic. We formalize the resulting marginal cluster support and establish simultaneous finite-sample concentration over all coordinates, explicit bounds for false inclusions and omissions, and exact support recovery when the minimum distributional separation dominates the high-dimensional stochastic error. We also quantify the dimension inflation induced by using an unadjusted testing level and give a familywise-error-controlled version. Under a conditional sufficiency condition, sure screening preserves the full-data posterior cluster probabilities, mutual information, and Bayes risk; an additional result characterizes robustness to imperfectly estimated cluster labels. The procedure is invariant to strictly increasing coordinate transformations and can retain low-variance cluster signals that principal components may discard. We further develop average dual information, a criterion combining partition agreement after transformation with structural coverage of cluster-relevant coordinates, and derive its basic properties and consistency. Simulations illustrate the theory, the interpretability of the selected coordinates, and the distinction between cluster-directed screening and variance-directed projection.
Sanoja Jha, Rishikesh Muralimohan, Praveen Athauda Arachchi et al.· 0 citations
Kernel spectral clustering with a single bandwidth can be inadequate for data exhibiting multiple characteristic pairwise-distance scales, a problem particularly prevalent in the high-dimensional regime. We address this issue through a multi-kernel formulation that aggregates kernels with different bandwidths. The bandwidths are selected as prescribed empirical quantiles of the pairwise squared distances, thereby capturing the relevant distance scales without requiring prior population-scale information. We develop a rigorous theoretical analysis of the resulting method under a general high-dimensional, multi-scale mixture model with heterogeneous cluster centers and covariance geometries. We construct a blockwise constant, low-rank informative approximation to the empirical multi-kernel matrix and establish row-wise $\ell_{2,\infty}$ perturbation bounds for its leading spectral components, as well as for the associated normalized Laplacian matrix. These bounds yield observation-level control of the spectral embedding, which is more informative than conventional global eigenspace perturbation estimates. Under suitable eigen-gap and cluster-separation conditions, we show that approximate $K$-means applied to the multi-kernel spectral embedding achieves exact recovery with high probability.
Zeqin Lin, Guangming Pan, Zhixiang Zhang et al.· 0 citations
Simulations show that VIMP and MPLOCO agree most closely when the fitted learner is well aligned with the data-generating mechanism, and clarify when intrinsic and extrinsic importance can be interpreted similarly and when they provide complementary information.
We introduce a robust nonparametric regression framework for functional covariates that combines functional principal component analysis (FPCA), marginal copula-scale normalization, bounded-score M-estimation, and multivariate Bernstein smoothing. The proposed procedure reduces the infinite-dimensional functional predictor to a low-dimensional score representation, transforms the retained scores onto the compact unit cube, and estimates a conditional M-functional through a smoothly aggregated system of local estimating equations. This construction is designed to accommodate nonlinear regression structure, heavy-tailed score distributions, and response contamination while limiting the influence of extreme observations. Under suitable regularity and undersmoothing conditions, we establish pointwise and uniform consistency, derive explicit convergence rates, and prove asymptotic normality. The limiting variance contains an explicit Bernstein concentration factor that plays a role analogous to the integrated squared kernel in classical nonparametric regression. The analysis also clarifies the interaction among the projection dimension, the Bernstein resolution, the empirical copula transformation, and the effective local sample size. The finite-sample performance of the method is examined through simulations involving heavy-tailed functional scores, Student-t errors, nonlinear regression effects, and increasing response contamination. The proposed estimator exhibits strong overall predictive performance and good robustness, with particularly favorable behavior under absolute-error criteria.
Wahiba Bouabsa, F. Alshahrani· Mathematics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.