AI Networking Cookbook: Practical recipes for AI-assisted network automation and development
Similar papers
Simultaneously Addressing Collinearity and Outliers: A Robust Two-Parameter Estimation Approach
Multicollinearity and outliers remain two major challenges in linear regression modeling, often occurring simultaneously in practical applications and leading to instability, inflated variance, and unreliable inference. Although shrinkage estimators such as ridge and Liu estimators effectively address multicollinearity, but are still sensitive to outliers. Conversely, robust estimators mitigate the influence of outliers but do not adequately resolve collinearity. This study proposes a new robust two-parameter shrinkage estimator (Rprop1) designed to simultaneously handle multicollinearity and outliers within the classical linear regression framework. The estimator is derived in canonical form, and its statistical properties are established through mean squared error (MSE) analysis. A comprehensive Monte Carlo simulation study is conducted across varying sample sizes, error variances, correlation levels, numbers of explanatory variables, and outlier magnitudes. The performance of the proposed estimator is compared with existing robust and shrinkage-based estimators, including robust ridge, robust Liu, and robust Kibra Lukman estimators. Simulation results consistently demonstrate that the proposed estimator achieves superior MSE performance across moderate to severe multicollinearity and outliers’ magnitudes and it’s also supported by the real life dataset. The findings suggest that the proposed method provides a more stable and efficient alternative for regression modeling in the presence of simultaneous collinearity and outliers.
Robust regression-based logarithmic estimators for efficient mean estimation in ranked set sampling
Accurate estimation of the population mean becomes challenging in the presence of outliers, especially when conventional estimators are implemented under simple random sampling. Ranked set sampling, known for its efficiency gains through judgment-based ranking, further suffers when extreme observations distort the estimation process. To address these issues, this study proposes a new class of robust regression-based logarithmic estimators for efficient population mean estimation under Ranked set sampling. The proposed estimators integrate robust regression techniques such as Huber-M, Huber-MM, least trimmed squares, least median of squares, Hampel-M, and Tukey-M with a logarithmic adjustment structure to minimize the impact of outliers while using auxiliary information. Theoretical properties such as bias and mean square error are derived under Ranked set sampling. Analytical efficiency comparisons demonstrate that the proposed class consistently achieves lower Mean sqaure error than several adapted families of robust estimators under Ranked set sampling. The theoretical results are validated with an extensive simulation study based on artificially generated symmetric and asymmetric populations and a real-data application with intentionally contaminated datasets, further validates the superiority of the proposed estimators.
A Simple Approximation to the Distribution of the Ridge Regression Estimator
We present a simple Gaussian approximation to the finite-sample distribution of the classical ridge regression estimator. Our approximation captures the fact that, in finite samples, the ridge regression estimator trades off bias and variance to reduce estimation and prediction error. Our approximation is based on nonstandard asymptotics where $i)$ we let the estimator's regularization parameter grow proportionally to the sample size; and $ii)$ we treat the population regression coefficients as \emph{local} to the reference vector that defines the estimator's direction of shrinkage. In contrast to other asymptotic approximations in the literature, we allow for general forms of heteroskedasticity and autocorrelation in the data generating process (at the cost of considering a low-dimensional model where the number of covariates is not allowed to grow with the sample size). We use our simple Gaussian approximation to propose two new strategies to select the regularization parameter for the ridge regression estimator. The suggested strategies select the regularization parameter to minimize either average or worst-case excess prediction risk, where risk is computed using our suggested Gaussian approximation.
Robust estimation of the autocorrelation function via forward ratios
It is obvious to say that an adequate estimation of the autocorrelation function is central in time series analysis. In this paper, we propose three new robust estimators based on ratios of observations, which offer strong resistance against outliers. While the first estimator, which is based on the median, is not efficient, the second is a Quasi Maximum Likelihood (QML) estimator with better efficiency properties. The third estimator is a plug-in estimator, which does not require numerical optimization and, consequently, is extremely simple from a computationally point of view, having similar efficiency to that of the ML estimator. We derive the asymptotic distribution of the first two estimators, when the true autocorrelations are zero. Furthermore, we also show that the asymptotic distribution of the plug-in estimator is rather close to that of the QML estimator, allowing for inference and, in particular, for the construction of point-wise significance bands for the autocorrelations. Using Monte Carlo simulations, we analyse the finite sample properties of the proposed estimators and compare them with those of the sample autocorrelations and alternative extant robust estimators based on ranks. Although the proposed estimators have larger dispersion than the sample autocorrelations in uncontaminated time series, they are highly robust in the presence of outliers. Also, they have better properties than popular alternative robust estimators based on ranks when estimating autocorrelations of order larger than one. The results are illustrated by estimating the correlogram of daily IBEX35 returns, quarterly US economic growth and monthly US inflation.
Generalised Regression Estimator for Population Variance in Simple Random Sampling
Accurate estimation of population variance plays a vital role in survey sampling, especially when simple random sampling is used. In this work, we propose a new generalized statistical inference in order to estimate the population variance using auxiliary information. We can use the relationship between the study variable and the auxiliary variable to construct a novel generalized class of estimators that is better performing in terms of minimum mean squared error (MSE) and has a higher percentage of relative efficiency than the traditional estimators. Theoretical properties of the proposed estimator such as the mean squared error, relative efficiency and bias are derived. The performance of the proposed generalized regression estimation estimator of the population variance under the simple random sampling design is assessed via simulation. The numerical findings reveal that the proposed estimator outperforms the competitors in all aspects. Also, the proposed estimator is robust as confirmed using various sample sizes and correlation coefficient. The research has made a significant contribution to the development of statistical procedures in survey sampling because the practical and efficient tools provided in the study were useful in estimating the variance.
A robust combined M-estimator for the negative binomial model
Robust inference for overdispersed count data is crucial in applications where outliers may substantially distort classical likelihood-based estimation of both the mean and dispersion. We develop robust estimation procedures for independent and identically distributed negative binomial data and provide practical guidelines on the choice of the estimator. We propose a combined robust M-estimator that jointly estimates the mean and dispersion parameter through an alternating updating scheme based on bounded score functions. The mean update relies on a bias-corrected Tukey-type M-estimator, extending the Poisson framework of Elsaied and Fried (2016) to the negative binomial setting, while the dispersion update follows a robust modification of score-based estimation in the spirit of Aeberhard et al. (2014). Under standard regularity conditions, we establish key theoretical guarantees for the proposed estimator, including consistency, local convergence of the alternating algorithm, asymptotic normality, and robustness to outliers through bounded influence. Extensive simulation studies compare the proposed method with maximum likelihood estimation, minimum disparity estimators, and weighted maximum likelihood approaches under both clean and contaminated sampling. The results show that classical likelihood procedures can be highly sensitive to additive contamination, particularly in dispersion estimation, whereas the proposed estimator achieves a favorable efficiency–robustness trade-off across a broad range of sample sizes and contamination regimes, while retaining high efficiency under the nominal model. An analysis of epileptic seizure count data further illustrates the practical relevance of the approach and the stability of the resulting inference without relying on ad hoc data cleaning.