Jul 2026· Statistical Methods in Medical Research· Vol 35, pp. 1658 - 1677· 0 citations· 46 references
Medicine
TL;DR
This paper compares several approaches for handling MNAR data in linear regression when missingness depends on both partially observed outcomes and predictors and concludes that not-at-random fully conditional specification was straightforward to implement and yielded coverage close to the nominal level in most scenarios.
Abstract
While most methods for missing not at random (MNAR) data in regression models address MNAR outcomes assuming fully observed predictors, real-world observational health and longitudinal studies often violate this assumption. This paper compares several approaches for handling MNAR data in linear regression when missingness depends on both partially observed outcomes and predictors. Through extensive simulations, we evaluate complete-case analysis, multiple imputation assuming missing at random, maximum likelihood estimation via the Heckman selection model, uncertainty intervals, multiple imputation under the Heckman selection model, not-at-random fully conditional specification, imputation stacking, and random indicator imputation. None of the methods consistently produced unbiased estimates or nominal coverage across all scenarios. However, not-at-random fully conditional specification was straightforward to implement and yielded coverage close to the nominal level in most scenarios, provided that the sensitivity parameters were specified near their true values. Our results highlight the importance of sensitivity analyses exploring various full-data models and careful parameter specification when addressing MNAR in both outcomes and predictors. We illustrate such sensitivity analyses using Betula study data on the relationship between longitudinal memory change and grey matter volume in aging. The association remained significant across most considered MNAR scenarios, reinforcing existing evidence for this relationship.
Missing data and confounding are common in real-world statistical applications, yet few studies have examined how imputation methods perform under time-varying confounding in binary variables, or how missingness mechanism, missing rate, missingness location and sample size jointly affect performance and the underlying identifiability conditions. We generated synthetic data and conducted a simulation study comparing missing data methods across scenarios varying these factors. Missingness was introduced in both treatment and outcome variables, and we applied stratified hot deck imputation, single mode imputation, multiple imputation with chained equations (MICE), and complete-case analysis. Average treatment effect (ATE) estimates were obtained using logistic regression with propensity score weighting, and we measured coverage, absolute bias and empirical standard errors across 48 scenarios with 500 replications each. Performance was primarily driven by the missingness mechanism and choice of method, with multiple imputation generally achieving better coverage and lower bias than other methods. Missingness location was also important, while missing rate and sample size primarily affected positivity violations, which were most pronounced under MNAR, high missingness and low sample sizes. Exchangeability violations from mild to moderate confounding were adequately controlled for by propensity score models, whereas strong confounding produced a modest decrease in coverage. Further research should examine additional ways identifiability conditions can be violated under missingness, using more advanced methods and more complex missingness scenarios.
B. Swallow, L. Brestrich, Victor Velasco-Pardo· 0 citations
For model comparison in random effects probit models with incompletely observed covariates, this paper develops a Bayesian data-augmentation workflow in which latent Gaussian responses, random effects, and missing covariate values are updated within a common augmented sampling scheme. Because specifying a fully parametric joint model for mixed continuous and categorical covariates is often unattractive in survey applications, missing covariates are updated by a decision-tree-assisted Bayesian-bootstrap step. Competing models are evaluated conditionally on one common medoid completion using Chib’s method with reduced Gibbs sampling; sensitivity is assessed with respect to the Chib evaluation point, the regression-coefficient prior, and the medoid reference model. The simulation study compares the proposed approach with complete case analysis, multiple imputation by chained equations, missForest single imputation, information criteria, and predictive criteria under MCAR, cross-dependent MAR-type, and self-masked MNAR scenarios. An empirical illustration based on the National Educational Panel Study demonstrates how the method can be used for comparing labor-market models of current employment when competence measures and employment-history covariates are incompletely observed. The results show that missing covariates can materially affect model rankings, and that the proposed workflow provides a transparent evidence-based comparison of nested and non-nested random effects probit specifications under incomplete covariate information.
Michael Bergrab· Statistical Methods & Ap...· 0 citations
In many studies, multiple longitudinal outcomes are collected, and interest lies in studying the association between these outcomes. Joint modeling is then required, but full likelihood estimation becomes infeasible as the number of outcomes increases. To address this, the pairwise-fitting approach was developed. However, the robustness of this pseudo-likelihood-based approach under missing at random (MAR) remains unclear. We investigate the impact of MAR dropout on the pairwise-fitting approach through a case and simulation study and compare the results to full likelihood estimation. In the simulation study, we simulate three continuous longitudinal outcomes so that full likelihood estimation remains computationally feasible, allowing a comparison with the pairwise fitting approach. Various settings are examined, including random intercept and random intercept-and-slope models, in which we vary the standard deviation of the error terms and the degree of correlation between random effects. Our results show that bias remains limited in random intercept models and in most random intercept-and-slope models. However, when the standard deviation of the error terms becomes large compared to that of the random effects, some bias appears in the covariances between the random effects of the outcomes not driving dropout. This bias is mitigated using multiple imputation. As a case study, we analyzed data from a schizophrenia study using both full likelihood and pseudo-likelihood approaches and compared the results.
Dries De Witte, G. Verbeke, T. Neyens et al.· Pharmaceutical statistics· 0 citations
Missing binary predictors are common in reliability, quality control, and industrial decision systems, yet imputation methods are often chosen by convenience rather than evidence. We conduct a Monte Carlo study comparing mode substitution, sequential hot‐deck, missForest, MICE, and KNN with three neighbourhood sizes under MCAR, MAR, and MNAR missingness, across missingness rates from 5% to 50% and two predictor‐dependence structures. Performance is evaluated on three targets: exact recovery of missing binary cells, recovery of logistic‐regression coefficients, and downstream classification using logistic regression, naive Bayes, support vector machines, and random forests. The results reveal a clear trade‐off. KNN is strongest for exact cell recovery under MCAR and MAR, whereas missForest performs best under MNAR. MICE is the most reliable choice for downstream predictive performance across learners and missingness mechanisms. By contrast, mode imputation and sequential hot‐deck achieve the best coefficient recovery. The main implication is operational: in binary‐data environments, imputation should be chosen to match the analytical objective–reconstruction, inference, or prediction–because no single method dominates all targets simultaneously.
Manuel Delfino, Fabio Rapallo· Quality and Reliability Engi...· 0 citations
Missing data in patient-reported outcome measures (PROMs) can affect estimation accuracy and inferential validity, particularly under missing-at-random (MAR) mechanisms. Multiple imputation (MI) is commonly recommended, but the relative performance of item-level and score-level imputation for longitudinal PROM composite scores remains insufficiently understood.
We conducted a simulation study based on longitudinal clinical trial PROM data. Predictive mean matching (PMM) and random forest (RF) imputation were evaluated at both the item and score level under empirically calibrated MAR mechanisms. Simulation scenarios varied sample size, overall missingness, and the proportion of unit nonresponse. Performance was assessed using root mean squared error (RMSE), bias magnitude, relative efficiency compared with complete case analysis (CCA), conditional confidence interval coverage, and variance ratio estimates.
RMSE decreased with increasing sample size and increased with higher proportions of missingness across all methods. Both PMM and RF generally outperformed CCA in terms of RMSE and relative efficiency, particularly under higher missingness. Differences between item-level and score-level imputation were generally modest, although score-level imputation tended to yield slightly lower RMSE and bias magnitudes under more severe missingness conditions. Conditional coverage remained close to nominal levels across most MI settings. Variance tended to be overestimated in some scenarios, although sensitivity analyses with larger numbers of imputations substantially reduced this effect. PMM showed comparatively stable performance across simulation settings.
Multiple imputation methods generally outperformed complete case analysis for handling longitudinal PROM data under MAR. While score-level imputation sometimes showed slightly more favorable performance than item-level imputation when targeting composite score means, the magnitude of these differences was frequently small in practical terms. PMM provided stable performance across a wide range of settings and represents a reasonable default approach for many applied PROM analyses.
Unknown authors· BMC Medical Research Methodo...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.