Skip to content
Open access

Simultaneously Addressing Collinearity and Outliers: A Robust Two-Parameter Estimation Approach

Aug 2026 · American Journal of Applied Statistics and Economics · 0 citations · 27 references

Abstract

Multicollinearity and outliers remain two major challenges in linear regression modeling, often occurring simultaneously in practical applications and leading to instability, inflated variance, and unreliable inference. Although shrinkage estimators such as ridge and Liu estimators effectively address multicollinearity, but are still sensitive to outliers. Conversely, robust estimators mitigate the influence of outliers but do not adequately resolve collinearity. This study proposes a new robust two-parameter shrinkage estimator (Rprop1) designed to simultaneously handle multicollinearity and outliers within the classical linear regression framework. The estimator is derived in canonical form, and its statistical properties are established through mean squared error (MSE) analysis. A comprehensive Monte Carlo simulation study is conducted across varying sample sizes, error variances, correlation levels, numbers of explanatory variables, and outlier magnitudes. The performance of the proposed estimator is compared with existing robust and shrinkage-based estimators, including robust ridge, robust Liu, and robust Kibra Lukman estimators. Simulation results consistently demonstrate that the proposed estimator achieves superior MSE performance across moderate to severe multicollinearity and outliers’ magnitudes and it’s also supported by the real life dataset. The findings suggest that the proposed method provides a more stable and efficient alternative for regression modeling in the presence of simultaneous collinearity and outliers.

Read PDF

Similar papers

Open access Sep 2026

Hybrid Principal Component-Based Estimators for Multicollinearity in Simultaneous Equation Models

This study addresses the problem of multicollinearity in simultaneous equation models (SEMs) by proposing hybrid estimators that integrate Principal Component Analysis (PCA) with conventional estimation techniques. Multicollinearity, characterized by high correlations among explanatory variables, adversely affects the efficiency and stability of estimators such as Two-Stage Least Squares (2SLS), Three-Stage Least Squares (3SLS), and Full Information Maximum Likelihood (FIML). To mitigate this problem, correlated regressors are transformed into orthogonal principal components prior to estimation, leading to the PCR-based SEM estimators. The performance of both classical and hybrid estimators is evaluated through Monte Carlo simulations under varying levels of multicollinearity, sample sizes and error structures. Mean Squared Error (MSE) is used as the evaluation criterion. The results indicate that the hybrid PCR-FIML estimator consistently outperforms competing estimators in most experimental settings. Overall, incorporating PCA into SEM estimation improves predictive accuracy, reduces estimation variance, and enhances robustness in the presence of multicollinearity.

A. Aladesuyi, O. Alabi, A. Bello · 0 citations
Open access Jul 2026

A New Efficient Biased Estimator in Overdispersed Count Regression for Effectively Handling Multicollinearity in COVID-19 Data from Saudi Arabia

Count data are widely encountered in many scientific fields, particularly in healthcare and epidemiology. One of the most commonly used approaches for analyzing such data is the negative binomial regression model (NBRM), due to its simplicity and effectiveness in modeling event frequencies. Despite its popularity, the presence of severe multicollinearity among explanatory variables can substantially inflate the variance of parameter estimates and reduce the reliability of statistical inference. To address this issue, this study proposes an improved shrinkage estimator for the NBRM, referred to as a novel class of negative binomial Liu-type estimator. The proposed estimator combines the advantages of ridge regression and the Liu estimator, aiming to reduce estimation variance while maintaining stable parameter estimates under conditions of multicollinearity. The proposed estimator is compared with the traditional maximum likelihood estimator, as well as existing ridge and Liu-type estimators, using performance measures such as the mean squared error. Its performance is evaluated through extensive Monte Carlo simulation experiments under different levels of multicollinearity and sample sizes. The simulation results demonstrate that the proposed estimator provides more accurate and stable estimates than the competing methods, particularly in the presence of high multicollinearity. To illustrate the practical applicability of the proposed approach, the method is applied to a real-world healthcare dataset related to COVID-19 cases in the Kingdom of Saudi Arabia. The empirical results confirm the effectiveness of the proposed estimator in improving estimation accuracy and model stability when modeling multivariate healthcare count data. Overall, the proposed estimator offers a useful alternative for modeling multicollinear healthcare count data and enhances the reliability of statistical analysis in applied health research.

E. H. Hafez, A. Hammad, R. Aldallal et al. · 0 citations
Open access Sep 2026

A robust combined M-estimator for the negative binomial model

Robust inference for overdispersed count data is crucial in applications where outliers may substantially distort classical likelihood-based estimation of both the mean and dispersion. We develop robust estimation procedures for independent and identically distributed negative binomial data and provide practical guidelines on the choice of the estimator. We propose a combined robust M-estimator that jointly estimates the mean and dispersion parameter through an alternating updating scheme based on bounded score functions. The mean update relies on a bias-corrected Tukey-type M-estimator, extending the Poisson framework of Elsaied and Fried (2016) to the negative binomial setting, while the dispersion update follows a robust modification of score-based estimation in the spirit of Aeberhard et al. (2014). Under standard regularity conditions, we establish key theoretical guarantees for the proposed estimator, including consistency, local convergence of the alternating algorithm, asymptotic normality, and robustness to outliers through bounded influence. Extensive simulation studies compare the proposed method with maximum likelihood estimation, minimum disparity estimators, and weighted maximum likelihood approaches under both clean and contaminated sampling. The results show that classical likelihood procedures can be highly sensitive to additive contamination, particularly in dispersion estimation, whereas the proposed estimator achieves a favorable efficiency–robustness trade-off across a broad range of sample sizes and contamination regimes, while retaining high efficiency under the nominal model. An analysis of epileptic seizure count data further illustrates the practical relevance of the approach and the stability of the resulting inference without relying on ad hoc data cleaning.

Hanan Elsaied, R. Fried · 0 citations
Open access Aug 2026

Robust regression-based logarithmic estimators for efficient mean estimation in ranked set sampling

Accurate estimation of the population mean becomes challenging in the presence of outliers, especially when conventional estimators are implemented under simple random sampling. Ranked set sampling, known for its efficiency gains through judgment-based ranking, further suffers when extreme observations distort the estimation process. To address these issues, this study proposes a new class of robust regression-based logarithmic estimators for efficient population mean estimation under Ranked set sampling. The proposed estimators integrate robust regression techniques such as Huber-M, Huber-MM, least trimmed squares, least median of squares, Hampel-M, and Tukey-M with a logarithmic adjustment structure to minimize the impact of outliers while using auxiliary information. Theoretical properties such as bias and mean square error are derived under Ranked set sampling. Analytical efficiency comparisons demonstrate that the proposed class consistently achieves lower Mean sqaure error than several adapted families of robust estimators under Ranked set sampling. The theoretical results are validated with an extensive simulation study based on artificially generated symmetric and asymmetric populations and a real-data application with intentionally contaminated datasets, further validates the superiority of the proposed estimators.

Renu Kumari, Anoop Kumar · 0 citations
Open access Sep 2026

A New Robust LQS-NTP Estimator for Mitigating Correlated Endogenous Variables and Extreme Observations in Linear Regression Models: Theoretical Development and Applications

This study proposes a Robust Least Quantile of Squares–New Two-Parameter (LQS-NTP) estimator for addressing multicollinearity and extreme observations in linear regression models. The proposed estimator combines the high-breakdown robustness of the Least Quantile of Squares (LQS) method with the shrinkage properties of the New Two-Parameter (NTP) estimator. The Mean Squared Error (MSE) of the proposed estimator was derived, and its biasing parameters were obtained by minimizing the corresponding MSE. The performance of the proposed estimator was assessed using real-life Boston Housing and Portland Cement datasets and compared with some already existing methods. The data sets exhibited substantial multicollinearity, together with several outlying observations. The proposed LQS-NTP estimator achieved the lowest MSE of both data employed. These results demonstrate that the proposed estimator provides an effective alternative for regression estimation in the simultaneous presence of multicollinearity and extreme observations.

Adewale Abdulahi Titilola, T. Olatayo, A. Taiwo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.