Jul 2026· Data Science in Science· Vol 5· 0 citations· 27 references
Medicine
Abstract
Missing data are common in environmental mixture studies and can bias inference if not properly addressed. This study evaluates how different imputation strategies influence variable selection performance in six mixture modeling frameworks: Weighted Quantile Sum (WQS), Bayesian WQS (BWQS), Quantile g-computation (Q-gcomp), Bayesian Kernel Machine Regression (BKMR), Elastic Net, and Least Absolute Shrinkage and Selection Operator (LASSO). Using Monte Carlo simulations (MCS), we generated multivariate normal (MVNORM) and multivariate t- (MVT) distributed exposures under linear and nonlinear outcome structures. Missingness was introduced at 25% for exposure and outcome variables, and at 5%–25% under both missing-at-random (MAR) and missing-not-at-random (MNAR) mechanisms. The study compares single imputation (SI) methods (mean, median, and k-nearest neighbors [KNN]) with multiple imputation approaches [MI] (MICE and Amelia) based on sensitivity (SE), specificity (SP), and false discovery rate (FDR). Across simulations, MI consistently improved variable selection performance under MAR, whereas SI and listwise deletion increased variability and reduced accuracy. Under MNAR, performance declined across all methods, with greater instability observed in flexible models such as BKMR and WQS. Q-gcomp demonstrated the most consistent balance across SE, SP, and FDR and remained relatively robust to violations of the MAR assumption. Additional analyses highlight the importance of imputation model specification: excluding relevant covariates (e.g. income) from the imputation process increased variability and reduced stability, particularly for flexible models and binary outcomes. Application to NHANES 2007–2014 data (n = 8,233) showed that MI improved the stability of variable importance estimates, with BWQS and Q-gcomp yielding more reproducible exposure rankings. Overall, results demonstrate that both the missingness mechanism and the choice and specification of imputation strategy substantially influence variable selection in mixture models, underscoring the importance of carefully designed imputation procedures in environmental epidemiology.
This paper compares several approaches for handling MNAR data in linear regression when missingness depends on both partially observed outcomes and predictors and concludes that not-at-random fully conditional specification was straightforward to implement and yielded coverage close to the nominal level in most scenarios.
T. Gorbach, Tim P. Morris, James E. Carpenter· Statistical Methods in Medic...· 0 citations
For model comparison in random effects probit models with incompletely observed covariates, this paper develops a Bayesian data-augmentation workflow in which latent Gaussian responses, random effects, and missing covariate values are updated within a common augmented sampling scheme. Because specifying a fully parametric joint model for mixed continuous and categorical covariates is often unattractive in survey applications, missing covariates are updated by a decision-tree-assisted Bayesian-bootstrap step. Competing models are evaluated conditionally on one common medoid completion using Chib’s method with reduced Gibbs sampling; sensitivity is assessed with respect to the Chib evaluation point, the regression-coefficient prior, and the medoid reference model. The simulation study compares the proposed approach with complete case analysis, multiple imputation by chained equations, missForest single imputation, information criteria, and predictive criteria under MCAR, cross-dependent MAR-type, and self-masked MNAR scenarios. An empirical illustration based on the National Educational Panel Study demonstrates how the method can be used for comparing labor-market models of current employment when competence measures and employment-history covariates are incompletely observed. The results show that missing covariates can materially affect model rankings, and that the proposed workflow provides a transparent evidence-based comparison of nested and non-nested random effects probit specifications under incomplete covariate information.
Michael Bergrab· Statistical Methods & Ap...· 0 citations
Missing data and confounding are common in real-world statistical applications, yet few studies have examined how imputation methods perform under time-varying confounding in binary variables, or how missingness mechanism, missing rate, missingness location and sample size jointly affect performance and the underlying identifiability conditions. We generated synthetic data and conducted a simulation study comparing missing data methods across scenarios varying these factors. Missingness was introduced in both treatment and outcome variables, and we applied stratified hot deck imputation, single mode imputation, multiple imputation with chained equations (MICE), and complete-case analysis. Average treatment effect (ATE) estimates were obtained using logistic regression with propensity score weighting, and we measured coverage, absolute bias and empirical standard errors across 48 scenarios with 500 replications each. Performance was primarily driven by the missingness mechanism and choice of method, with multiple imputation generally achieving better coverage and lower bias than other methods. Missingness location was also important, while missing rate and sample size primarily affected positivity violations, which were most pronounced under MNAR, high missingness and low sample sizes. Exchangeability violations from mild to moderate confounding were adequately controlled for by propensity score models, whereas strong confounding produced a modest decrease in coverage. Further research should examine additional ways identifiability conditions can be violated under missingness, using more advanced methods and more complex missingness scenarios.
B. Swallow, L. Brestrich, Victor Velasco-Pardo· 0 citations
Missing binary predictors are common in reliability, quality control, and industrial decision systems, yet imputation methods are often chosen by convenience rather than evidence. We conduct a Monte Carlo study comparing mode substitution, sequential hot‐deck, missForest, MICE, and KNN with three neighbourhood sizes under MCAR, MAR, and MNAR missingness, across missingness rates from 5% to 50% and two predictor‐dependence structures. Performance is evaluated on three targets: exact recovery of missing binary cells, recovery of logistic‐regression coefficients, and downstream classification using logistic regression, naive Bayes, support vector machines, and random forests. The results reveal a clear trade‐off. KNN is strongest for exact cell recovery under MCAR and MAR, whereas missForest performs best under MNAR. MICE is the most reliable choice for downstream predictive performance across learners and missingness mechanisms. By contrast, mode imputation and sequential hot‐deck achieve the best coefficient recovery. The main implication is operational: in binary‐data environments, imputation should be chosen to match the analytical objective–reconstruction, inference, or prediction–because no single method dominates all targets simultaneously.
Manuel Delfino, Fabio Rapallo· Quality and Reliability Engi...· 0 citations
Missing values, atypical observations, and heterogeneity across latent groups are common sources of complexity in regression data. The contaminated Gaussian cluster-weighted model (CG-CWM) provides a natural framework for handling atypical observations, including outliers and leverage points, in model-based clustering. We extend the CG-CWM to data with missing-at-random (MAR) values in both the response and covariate spaces. The proposed model provides clustering in regression analysis while distinguishing typical observations, outliers, and good and bad leverage points. By treating covariates as random, the model preserves assignment dependence, allowing them to contribute directly to cluster formation. Maximum likelihood estimation is performed through an expectation-conditional maximization (ECM) algorithm that accounts for four sources of incomplete information: missing responses and covariates, unknown component memberships, and latent contamination indicators. Conditional on these indicators, the joint distribution of responses and covariates is multivariate Gaussian, yielding closed-form conditional distributions for missing values and incorporating missingness uncertainty directly into parameter updates. Thus, missing values are handled within model fitting rather than by preliminary imputation. The framework provides clustering, clusterwise regression, model-based treatment of MAR values, and detection of atypical observations. Performance is assessed through numerical studies under varying levels of contamination and missingness patterns, and a real data application.
Missing Health-Related Quality of Life (HRQoL) data in clinical studies risk propagating bias into health technology assessments (HTAs) and cost-utility analyses. Despite this, current National Institute for Health and Care Excellence (NICE) guidance offers no specific recommendations for handling missing HRQoL values. Using Monte Carlo simulations (1000 datasets), this study evaluated nine imputation methods, across the missing completely at random (MCAR), missing at random (MAR) and missing not at random (MNAR) assumptions at levels ranging from 5% to 50%. Performance was assessed using bias, variance, and coverage of the true HRQoL mean. While multiple imputation by chained equations (MICE)-based approaches performed best under MCAR and MAR, all methods showed bias under MNAR, with a delta-pattern mixture model performing the best (relative bias ≤2.1% at all missingness levels but coverage falls to 35.4% at 50% missingness). The choice of imputation method is key to preventing biased results from propagating into cost-effectiveness analysis, which in theory may lead to suboptimal reimbursement decisions and inefficient healthcare spending. To address the lack of explicit guidance from HTA bodies, we have developed a preliminary policy that could be used for HTA submissions: If the missingness pattern is not known and missingness ≤5%, it is suggested that most methods (except GLM) are acceptable (though care should be taken when using CCA and LOCF if MNAR is suspected). It is also suggested that MICE-based techniques are used as the base case for missingness >5%, and delta-PMM used as a sensitivity analysis when data are not MCAR. Further simulation studies would be required to strengthen the suggestions in this preliminary policy; however, the development of universal recommendations would lead to improved consistency and reliability of HTA globally.
J. Moss, N. Hansell, Erin Barker et al.· Health Economics and Policy· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.