Aug 2026· Statistical Methods & Applications· 0 citations· 68 references
Abstract
For model comparison in random effects probit models with incompletely observed covariates, this paper develops a Bayesian data-augmentation workflow in which latent Gaussian responses, random effects, and missing covariate values are updated within a common augmented sampling scheme. Because specifying a fully parametric joint model for mixed continuous and categorical covariates is often unattractive in survey applications, missing covariates are updated by a decision-tree-assisted Bayesian-bootstrap step. Competing models are evaluated conditionally on one common medoid completion using Chib’s method with reduced Gibbs sampling; sensitivity is assessed with respect to the Chib evaluation point, the regression-coefficient prior, and the medoid reference model. The simulation study compares the proposed approach with complete case analysis, multiple imputation by chained equations, missForest single imputation, information criteria, and predictive criteria under MCAR, cross-dependent MAR-type, and self-masked MNAR scenarios. An empirical illustration based on the National Educational Panel Study demonstrates how the method can be used for comparing labor-market models of current employment when competence measures and employment-history covariates are incompletely observed. The results show that missing covariates can materially affect model rankings, and that the proposed workflow provides a transparent evidence-based comparison of nested and non-nested random effects probit specifications under incomplete covariate information.
This paper compares several approaches for handling MNAR data in linear regression when missingness depends on both partially observed outcomes and predictors and concludes that not-at-random fully conditional specification was straightforward to implement and yielded coverage close to the nominal level in most scenarios.
T. Gorbach, Tim P. Morris, James E. Carpenter· Statistical Methods in Medic...· 0 citations
Missing data are common in environmental mixture studies and can bias inference if not properly addressed. This study evaluates how different imputation strategies influence variable selection performance in six mixture modeling frameworks: Weighted Quantile Sum (WQS), Bayesian WQS (BWQS), Quantile g-computation (Q-gcomp), Bayesian Kernel Machine Regression (BKMR), Elastic Net, and Least Absolute Shrinkage and Selection Operator (LASSO). Using Monte Carlo simulations (MCS), we generated multivariate normal (MVNORM) and multivariate t- (MVT) distributed exposures under linear and nonlinear outcome structures. Missingness was introduced at 25% for exposure and outcome variables, and at 5%–25% under both missing-at-random (MAR) and missing-not-at-random (MNAR) mechanisms. The study compares single imputation (SI) methods (mean, median, and k-nearest neighbors [KNN]) with multiple imputation approaches [MI] (MICE and Amelia) based on sensitivity (SE), specificity (SP), and false discovery rate (FDR). Across simulations, MI consistently improved variable selection performance under MAR, whereas SI and listwise deletion increased variability and reduced accuracy. Under MNAR, performance declined across all methods, with greater instability observed in flexible models such as BKMR and WQS. Q-gcomp demonstrated the most consistent balance across SE, SP, and FDR and remained relatively robust to violations of the MAR assumption. Additional analyses highlight the importance of imputation model specification: excluding relevant covariates (e.g. income) from the imputation process increased variability and reduced stability, particularly for flexible models and binary outcomes. Application to NHANES 2007–2014 data (n = 8,233) showed that MI improved the stability of variable importance estimates, with BWQS and Q-gcomp yielding more reproducible exposure rankings. Overall, results demonstrate that both the missingness mechanism and the choice and specification of imputation strategy substantially influence variable selection in mixture models, underscoring the importance of carefully designed imputation procedures in environmental epidemiology.
Y. S. Boafo, S. Mostafa, E. Obeng-Gyasi· Data Science in Science· 0 citations
Longitudinal data modeling attracts special interest from researchers and practitioners because of its capacity to help understand individual and mean trends of growth and development over time. One feature of longitudinal design is the collection of extensive information on participants to help explain the underlying growth process. However, selecting which covariates or independent variables (i.e., predictors) to include in a statistical model is challenging especially with the goals of avoiding both overfitting and underfitting. The present study demonstrates and compares multiple Bayesian variable selection methods for the selection of covariates to explain longitudinal growth processes, including both shrinkage methods (e.g., horseshoe) and stochastic variable search method (e.g., spike-and-slab priors and their extension to a Normal Mixture of Inverse Gammas), via an extensive Monte Carlo simulation study using a piecewise random effects model with unknown change point as an example. Our goal is to provide recommendations for researchers and practitioners about the advantages and limitations of Bayesian variable selection methods for variable selection in longitudinal studies, including model convergence and parameter estimation precision and accuracy.
Yue Zhao, Nidhi Kohli, Eric F. Lock· Multivariate Behavioral Rese...· 0 citations
Analyzing zero-inflated count data in single-case experimental designs presents a significant analytical challenge. Although traditional zero-inflated generalized linear mixed models are available, they estimate a conditional treatment effect given an individual coming from the count process, which often mismatches the applied researcher's interest in the overall effect of an intervention for each subject. These conditional models also face interpretational and estimation challenges within the small-sample context of single-case experimental designs. This study introduces and evaluates a Bayesian marginalized zero-inflated Poisson (mZIP) model with random effects. This framework reparameterizes the model to estimate the marginal intervention effect, aligning the statistical estimand with the typical research question. A Monte Carlo simulation study was conducted to evaluate the mZIP model's performance. We compared its performance with other methods that also target the marginal treatment effect: the log response ratio (LRR), a Poisson generalized linear mixed-effects model (GLMM), and a negative binomial GLMM. Simulation results indicate that the Bayesian mZIP model consistently recovers unbiased estimates of the marginal treatment effect and provides reliable statistical inference. The LRR, Poisson GLMM, and negative binomial GLMM produced biased estimates for the marginal effect under conditions with smallest sample size and highest zero-inflation rate. They also suffered from low coverage rate, making their inferential statistics invalid. The LRR indicated lower statistical power than the mZIP model. We apply the mZIP model to two empirical data sets to illustrate its application and interpretation. Finally, we discuss the distinction between conditional and marginal estimands, as well as the limitations and future directions. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Trial-based economic evaluations are widely used to assess the cost-effectiveness of healthcare interventions and inform decision-making. Cost and effectiveness outcomes are typically collected using multi-item questionnaires administered at multiple time points, and are often subject to item-level missingness. In principle, imputation (i.e., replacing missing value with estimated or substituted values) should be performed at the item level to fully exploit available information. However, this is rarely implemented in practice due to several statistical challenges, including the longitudinal data structure, cross-item dependence, heterogeneous missingness patterns, and the mixture of skewed cost and count data. In this paper, we develop a Bayesian longitudinal model for imputing item-level missing data in trial-based economic evaluations that accommodates these complexities within a unified framework. The approach combines a transition-model formulation for longitudinal dependence with flexible distributional assumptions and explicit modelling of cross-item relationships, allowing item-level responses of different types to be coherently modelled over time. Motivated by a real-world trial, we demonstrate the flexibility and practical applicability of the proposed approach. We further discuss how the model can be extended to settings where data may be missing not at random.
Xiaoxiao Ling, A. Gabrio, Gianluca Baio· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.