Skip to content
Open access

Bootstrap-enhanced regularization addressing multicollinearity and skewness in high-dimensional immunophenotyping data

Aug 2026 · BMC Bioinformatics · 0 citations

TL;DR

The Bootstrap-Enhanced Regularization Method (BERM) is a robust approach for variable selection and coefficient estimation in complex biomedical datasets, achieving the highest overall balanced accuracy while maintaining competitive coefficient estimation performance across a range of simulated sparsity, noise, and dimensionality scenarios.

Abstract

Accurate identification and estimation of variables associated with outcomes or disease states are critical for advancing diagnosis, prognosis, and precision medicine in biomedical research. Regularized regression techniques, such as lasso, are widely employed to enhance interpretability by reducing model complexity and identifying significant variables. However, these methods face two major challenges: (1) the exclusion of important variables due to high correlation with included predictors, and (2) the presence of skewness in human biomedical datasets, which violates key statistical assumptions. Current approaches that fail to address these issues simultaneously may lead to biased interpretations and unreliable coefficient estimates. To overcome these limitations, we propose an enhanced two-step approach, the Bootstrap-Enhanced Regularization Method (BERM). BERM outperformed existing regularization methods in variable selection, achieving the highest overall balanced accuracy while maintaining competitive coefficient estimation performance across a range of simulated sparsity, noise, and dimensionality scenarios. We further demonstrated the effectiveness of BERM by applying it to a human immunophenotyping dataset to identify important immune parameters in the autoimmune disease, type 1 diabetes. BERM is a robust approach for variable selection and coefficient estimation in complex biomedical datasets. Its consistent performance across a wide range of data conditions supports more reliable identification of important variables. An open-source implementation of BERM is available as an R package on GitHub ( https://github.com/xiaorudong/berm ).

Read PDF

Similar papers

Open access Jul 2026

Revisiting Logistic Regression for High-Dimensional Gene Expression Data

Logistic regression remains a widely used classification method due to its interpretability and computational efficiency, but its direct application to high-dimensional biomedical data is limited when the number of features greatly exceeds the number of samples. In this paper, we propose a reformulated logistic regression framework designed for feature selection and classification in complex high-dimensional settings. The method is evaluated on three biomedical datasets, including scenarios with tens of thousands of attributes and substantially fewer samples. Across these datasets, the proposed approach achieved clear separation between control and disease groups while selecting a compact set of features. Several selected features were consistent with previously reported disease-associated markers, supporting the biological plausibility of the model, while additional selected features suggest potential novel candidates for further investigation. These results indicate that the proposed framework may provide an interpretable and computationally efficient alternative for feature selection in high-dimensional computational biology applications.

Rossana O. Souza, W. Rodrigues, Bráulio Couto et al. · 0 citations
Open access Aug 2026

A simulation-based comparison of Boruta, LASSO, and Elastic Net for variable selection in logistic regression, with an ovarian cancer miRNA application

Variable selection is a central challenge in logistic regression, particularly in high-dimensional biomedical applications where correlated predictors and limited sample sizes complicate reliable identification of relevant variables. This study aims to systematically compare three widely used variable selection approaches - Boruta, LASSO, and Elastic Net - under a range of data-generating conditions and to illustrate their performance using an ovarian cancer miRNA dataset. We conducted a simulation study across 36 logistic regression scenarios varying in sample size, dimensionality, predictor correlation, and effect magnitude. Performance was evaluated using true-positive and false-positive selection rates. In addition, all three methods were applied to a real-world serum miRNA expression dataset, and discriminative performance was assessed using the area under the receiver operating characteristic curve (AUC). Boruta, Elastic Net, and LASSO exhibited distinct variable selection behaviors across simulation scenarios. Boruta maintained strong true-positive recovery while controlling false positives in most settings, particularly when predictors were highly correlated. Elastic Net consistently achieved high sensitivity but produced comparatively large false-positive rates. LASSO showed the most conservative behavior, recovering fewer true predictors while maintaining low false-positive rates across nearly all scenarios. In the ovarian cancer miRNA application, all three methods achieved similarly strong test-set AUC performance, despite marked differences in the size of the selected biomarker panels. The results demonstrate clear trade-offs among the three methods. Boruta offers a favorable balance between sensitivity and specificity in highly correlated settings. Elastic Net prioritizes sensitivity at the cost of increased false discoveries, whereas LASSO provides stricter false-positive control with reduced sensitivity. These findings offer practical guidance for selecting variable selection methods in logistic regression, particularly for high-dimensional biomedical applications.

Reza Arabi Belaghi, Hulya Yurekli, Farzaneh Hamidi et al. · 0 citations
Aug 2026

Beyond predictive performance: Interpretability challenges and feature importance bias in XGBoost-based readmission models.

Song et al. report a machine-learning framework based on the eXtreme Gradient Boosting (XGBoost) algorithm for predicting 1-year unplanned readmissions among elderly patients with coronary heart disease (CHD). This commentary examines critical limitations in the interpretability and methodological robustness of such models. Extensive prior work has demonstrated that tree‑based algorithms can exhibit structural biases in feature‑importance estimates, particularly when predictors display substantial collinearity or heterogeneous measurement scales. SHAP explanations, by inheriting these model‑embedded biases, may overstate the relevance of variables whose prominence arises from algorithmic artifacts rather than clinically coherent patterns. To enhance reliability, model‑agnostic validation, non‑parametric association analyses, and unsupervised strategies that mitigate multicollinearity should complement predictive modeling efforts. Strengthening interpretability frameworks is essential to ensure that machine‑learning-derived insights meaningfully inform cardiovascular care and support reproducible clinical translation.

S. Oka, Maito Suzuki, Y. Takefuji · 0 citations
Open access Jul 2026

Stability and interpretability of penalized logistic regression models for breast cancer risk prediction

Penalized logistic regression is widely used in biomedical classification to address multicollinearity and improve predictive performance, yet the stability and reproducibility of selected predictors are often overlooked. This study evaluates feature stability and interpretability in ridge, lasso, and elastic-net logistic regression for breast cancer diagnosis using the Wisconsin Diagnostic Breast Cancer dataset. Models were trained with cross-validated tuning and evaluated on an independent test set using discrimination, classification, and calibration metrics. Feature stability was quantified through bootstrap selection frequencies. All penalized models achieved near-perfect discrimination and improved calibration compared with unpenalized logistic regression. However, substantial differences emerged in stability and sparsity. Ridge regression exhibited maximal stability but retained all predictors, limiting interpretability. Lasso regression produced highly sparse models but showed greater selection variability. Elastic-net regression balanced sparsity and stability, consistently retaining correlated predictors linked to tumor morphology. These findings demonstrate that stability assessment provides critical information beyond predictive accuracy and supports stability-aware penalized modeling for interpretable and reproducible biomedical risk prediction.

F. Okyere, Michael Nyanney · 0 citations
Jul 2026

Assessing penalized approaches for estimating causal treatment effects under extremely limited overlap in oncology.

BACKGROUND Overlap weighting (OW) is increasingly used to estimate treatment effects in observational cancer studies. OW has attractive features: it targets the clinical equipoise population and mitigates the influence of extreme propensity score (PS) weights. Additionally, under regularity conditions, when the PS model is fitted using a standard logistic regression model (LRM) with all observed covariates included, OW achieves exact covariate balance between treated and control groups, meaning that the standardized mean differences for the included covariates are zero. However, in oncology data, imbalanced treatment patterns, small subgroups, and limited PS overlap frequently cause separation and non-convergence, undermining stable estimation in logistic regression. METHODS We evaluated five methods for constructing PSs for estimating the average treatment effect in the overlap population (ATO): standard logistic regression, Firth's penalized logistic regression, a double-penalized logistic regression method that combines Firth's correction with ridge regularization, and two variants designed to preserve exact covariate balance. Performance was assessed through Monte Carlo simulations under varying overlap, data complexity, and model misspecification. We also applied these methods to a retrospective cohort of 5348 patients with early-stage breast adenocarcinoma and Charlson Comorbidity Index ≥2 from the multi-institutionally linked nationwide data to estimate the effect of definitive surgery on 3-year all-cause mortality. RESULTS In simulations, standard LRMs showed unstable estimation or non-convergence in finite-sample settings characterized by low treatment prevalence and limited effective overlap. Penalized methods improved numerical stability, reduced extreme PS values, and generally showed better finite-sample performance, particularly when the LRM is not converged. In the breast cancer study, only 2.1% of patients did not undergo surgery, indicating marked treatment imbalance. Overall estimates were similar across methods, but in patients aged <40 years, a LRM yielded an extreme ATO estimate, whereas Firth's and double-penalized methods produced more stable and consistent results. CONCLUSIONS In oncology subgroups where low treatment prevalence and limited effective overlap lead to unstable or non-convergent LRMs, penalized regression approaches provide a practical strategy for improving ATO estimation.

Sangwon Lee, H. Chae, Dong-woo Choi et al. · 0 citations
Open access Aug 2026

metadeconfoundR: covariate analysis of high-dimensional cross-sectional omics data

Identifying disease biomarkers from large molecular datasets is complicated by correlated and confounded signals like comorbidities and treatment regimens, batch effects, and cohort biases. These effects bias statistical inference and clinical conclusions. Robust methodologies are fundamental for reliable biomarker discovery. metadeconfoundR is an R package for conservative biomarker discovery in (multi-)omics case-control datasets. It has a scalable two-step confounder-aware statistical framework for retaining only associations with independent support. It identifies covariate-naive univariate associations between omics features and metadata, then re-evaluates these associations using parallel post-hoc nested linear model testing to account for potential confounders. Confounded associations are flagged if they fully reduce to at least one other variable. metadeconfoundR supports parallel computation for large-scale datasets, offers visualization and tools for interpreting results and secondary analyses. We benchmark metadeconfoundR against state-of-the-art methods for identifying biomarkers using simulated ground truth derived from microbiome data, and demonstrate its ability to disentangle confounding effects while preserving statistical power, offering particular advantage when multiple covariates are present. metadeconfoundR functions for any -omics data type with continuous or categorical metadata/covariates. metadeconfoundR is available on CRAN (https://cran.r-project.org/web/packages/metadeconfoundR/) and GitHub (https://github.com/TillBirkner/metadeconfoundR).

T. Birkner, Chia-Yu Chen, Morgan Essex et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.