Skip to content
Preprint

A directional Hosmer-Lemeshow goodness-of-fit test for sparse logistic regression

Jul 2026 · 1 citation · 25 references
Mathematics

Abstract

Goodness-of-fit assessment for the binary logistic regression model is difficult when covariates are continuous: the data are effectively sparse, the classical Pearson and deviance tests fail, and practitioners rely on partition-based tests, such as the Hosmer-Lemeshow test, that group observations before comparing observed and expected counts. We study a partition test that modifies the Hosmer-Lemeshow statistic with a single directional correction term, weighted by $(1-2\bar\pi_g)$ and referred to a $\chi^2_{G-2}$ distribution. The correction is the grouped form of the Osius-Rojek/Farrington standardization; grouping makes it well defined in the sparse regime, and it targets the asymmetric over- and under-prediction that a misspecified link induces. A single alignment functional captures its effect, predicting where the test gains power (asymmetric-link misspecification) and where it does not (symmetric departures, and covariate-space structure that no probability-grouping test can see). In simulations the test holds its size; no well-calibrated partition test is more sensitive to asymmetric-link misfit, and it clearly exceeds Hosmer-Lemeshow there, most so for the complementary log-log link -- a modest gain that fades as $n$ grows; it ties Hosmer-Lemeshow on an omitted interaction and is less powerful on an omitted quadratic (by about ten percentage points at $n=1000$). A real-data application illustrates its use, and the test is implemented in the R package ebrahim.gof.

View source

Similar papers

Preprint Sep 2026

Shrinkage invalidates the Hosmer-Lemeshow test: goodness of fit for penalized logistic regression, with an application to glaucoma diagnosis

Clinical prediction models are increasingly fitted by penalized logistic regression, because collinearity or many candidate predictors makes maximum likelihood unstable or impossible. Calibration is then almost always assessed by a grouped goodness-of-fit test such as the Hosmer-Lemeshow test. We show that this combination is invalid. Under ridge regression the grouped standardized residuals acquire a non-centrality induced by shrinkage, so the reference distribution used in practice is wrong, and at the penalty that most improves the fitted probabilities the test rejects correctly specified models between 92 and 100 per cent of the time. We derive the corrected law and define the shrinkage-corrected Hosmer-Lemeshow test, which subtracts an estimate of that non-centrality, restoring the maximum likelihood reference exactly to first order, and is made valid by prepivoting at a power cost we measure. We also give the attenuation law governing what any such test can detect once the linear predictor must be estimated. In glaucoma diagnosis by confocal laser tomography, where the maximum likelihood estimate does not exist, the corrected test finds the evidence for misfit weaker by more than three orders of magnitude: the fitted risks are too flat rather than mis-ordered, so the model needs recalibration rather than rebuilding.

Unknown authors · 0 citations
Preprint Aug 2026

Goodness-of-Fit Tests and Calibration Machine-Learning Algorithms for Logistic Regression with Sparse Data

Assessing the goodness-of-fit of a logistic regression model is a critical prerequisite before the model is used for inference. However, goodness-of-fit (GOF) tests such as the chi-square and deviance tests often give invalid results when the data are"sparse"-- a common issue with continuous predictors like age or weight, where the asymptotic distributional assumptions are not satisfied. This thesis studies classical GOF tests for binary logistic regression under both grouped and sparse data, comparing about 30 statistical tests and machine-learning calibration algorithms. These span the classical chi-square and Hosmer-Lemeshow variants, standardized Pearson statistics, covariate-space partitioning, smoothing-based methods, and contemporary calibration machine-learning and bootstrap procedures. At a fixed size, the GiViTI calibration test (2016), McCullagh (1989), Osius-Rojek (1992), le Cessie (1995) and Stute-Zhu (2002) proved empirically powerful, balancing correct identification of bad models (high empirical power) against not raising false alarms on good models (correct empirical Type I error). Relying on formal methods alone is insufficient: visual diagnostics such as calibration plots are a vital exploratory step for detecting model deficiencies that formal tests often overlook. An application to real data (the Low Birth Weight dataset) shows that many of these tests fail to give valid conclusions when exposed to the complexities of actual datasets. The main conclusion is that model assessment requires a combination of several powerful statistical tests alongside careful visual inspection of model calibration.

Ebrahim Khaled Ebrahim · 0 citations
Preprint Jul 2026

Benchmarking Goodness-of-Fit and Calibration Algorithms for Logistic Regression Classifiers: A Large-Scale Simulation Study under Sparse Data

This paper provides a unified taxonomy and a large-scale, reproducible simulation benchmark; more than twenty tests are implemented in the open-source R package ebrahim.gof, and translate these findings into practical, evidence-based guidance for assessing logistic regression fit.

Ebrahim Khaled Ebrahim, Ahmed El-Kotory · 0 citations
Preprint Jul 2026

Partial pooling predicts cross-validation reliability: a closed-form triage and Rao-Blackwellised cure for hierarchical LOO

For hierarchical models, Pareto-smoothed importance-sampling leave-one-out cross-validation (PSIS-LOO) fails on the folds where a random-effect coordinate is data-driven and its group is small. We show that the Gelman-Pardoe pooling factor and structural leverage predict these folds from model structure and group sizes, without forming importance weights. In Gaussian linear mixed models the leverage reduces to group size, giving a design-time map that separates the failing ($\hat{k}>0.7$) folds with AUC 0.96; across replicated logistic GLMMs the post-fit, weight-free predictor reaches AUC 0.81. The cure is integrated importance sampling: marginalise the random-effect block and importance-sample only the base parameters. This is not new, but we contribute its observation-level specialisation for random-intercept GLMMs: an analytic Gaussian downdate and a 1-D quadrature for Bernoulli, binomial and Poisson responses, packaged as a drop-in rb_loo(fit). Against exact refits, this marginalised estimator (RB-LOO) is $3\times$ more accurate than moment matching on singleton-heavy logistic GLMMs. On overdispersed count data with 97 failing folds, moment matching leaves 37 uncorrected and is no more accurate than raw PSIS-LOO, while RB-LOO reproduces the 82-minute exact refit (elpd RMSE 0.04) at no cost. The error changes decisions: against a negative-binomial model, PSIS-LOO reports decisive evidence ($z=4.9$) and reloo reports significant evidence ($z=3.4$) for the more complex model, where an exact analysis, reproduced by RB-LOO, finds the two indistinguishable ($z=1.0$). A base-fiber Schur decomposition splits case-deletion influence into a vertical (pooling) term that governs where PSIS-LOO fails and a horizontal (variance-component) term that governs where RB-LOO is itself strained, giving a two-level triage that recovers the exact answer while refitting only the few folds that need it.

Aidan D. Bindoff · 0 citations
Preprint Aug 2026

Distribution-free testing of linear type

We introduce a distribution-free goodness-of-fit test, termed the omega-1 test, which naturally complements the Kolmogorov--Smirnov test and Cram\'{e}r--von Mises test and can be viewed as their (piecewise) linear analog. Defined as an $\mathrm{L}^{1}$-functional of the empirical process, the test statistic improves on balancing sensitivity to localized and diffuse alternatives and gives a robust and interpretable measure of distributional discrepancy, apart from close connections to the Wasserstein 1-distance. For finite samples, we derive a finite-dimensional computational form for the statistic under general conditions, which leads to various explicit formulas for its null distribution. Under mild continuity assumptions, the limiting statistic is distribution-free, with explicit distribution formulas. In composite settings, the statistic is also compatible with the Khmaladze transformation, enabling asymptotically distribution-free testing. The limiting transformed statistic also has an explicit distribution that escapes reliance on intractable compensator processes or purely numerical evaluation. Simulation results indicate rapid convergence of the finite-sample distributions to their limiting counterparts and support the practical applicability of the test.

Wei-Xu Xia · 0 citations
Preprint Jul 2026

Testing for correct model specification in copula regression models

We propose a goodness-of-fit test for semiparametric copula regression models. Such models express the regression function in terms of marginal distribution functions and copula densities and therefore provide a flexible way to avoid fully nonparametric estimation in high-dimensional regression problems. Their performance, however, depends crucially on the specification of the parametric copula family. Instead of testing the copula model itself, we assess misspecification directly at the level of the induced regression function. To this end, we introduce a weighted $L^2$-distance between the true regression function and its best approximation within the postulated copula regression model. A kernel-based estimator of this distance is proposed and shown to be consistent and asymptotically normal under both the null hypothesis of correct specification and fixed alternatives. We derive a classical specification test and, using a self-normalized sequential statistic, construct pivotal confidence intervals and tests for relevant deviations from the model. Finite-sample simulations demonstrate accurate level approximation and good power properties of the proposed procedures.

Holger Dette, Philipp Dörr · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.