Skip to content
Open access

Columnwise neural imputation for incomplete ordinal psychometric data.

Aug 2026 · Psychological methods · 0 citations
Medicine

TL;DR

This work proposes the columnwise neural imputation (COLNI) algorithm, a artificial neural network-based imputation approach to impute missing ordinal responses in psychometric data, and evaluates its effectiveness in a multidimensional empirical setting.

Abstract

Missing data are pervasive in psychological and educational assessments. Naive remedies, including listwise deletion and item-mean imputation, often degrade research validity and misinform subsequent decisions. Recent advances in artificial neural networks have demonstrated their efficacy in prediction-related tasks by using observed features to infer unknown values. Building on this potential, we propose the columnwise neural imputation (COLNI) algorithm to impute missing ordinal responses in psychometric data. Simulation studies demonstrated that, when benchmarked against conventional methods, COLNI more accurately recovered item means, inter-item correlations, and person and item parameters under the multidimensional graded response model. We further evaluated COLNI using data from the Short Dark Triad test, confirming its effectiveness in a multidimensional empirical setting. We conclude with implementation guidelines and avenues for refining and extending this artificial neural network-based imputation approach in future research. (PsycInfo Database Record (c) 2026 APA, all rights reserved).

Read PDF

Similar papers

Open access Jul 2026

Not All Missing Data are Equal: Choosing the Right Imputation Method for Binary Datasets

Missing binary predictors are common in reliability, quality control, and industrial decision systems, yet imputation methods are often chosen by convenience rather than evidence. We conduct a Monte Carlo study comparing mode substitution, sequential hot‐deck, missForest, MICE, and KNN with three neighbourhood sizes under MCAR, MAR, and MNAR missingness, across missingness rates from 5% to 50% and two predictor‐dependence structures. Performance is evaluated on three targets: exact recovery of missing binary cells, recovery of logistic‐regression coefficients, and downstream classification using logistic regression, naive Bayes, support vector machines, and random forests. The results reveal a clear trade‐off. KNN is strongest for exact cell recovery under MCAR and MAR, whereas missForest performs best under MNAR. MICE is the most reliable choice for downstream predictive performance across learners and missingness mechanisms. By contrast, mode imputation and sequential hot‐deck achieve the best coefficient recovery. The main implication is operational: in binary‐data environments, imputation should be chosen to match the analytical objective–reconstruction, inference, or prediction–because no single method dominates all targets simultaneously.

Manuel Delfino, Fabio Rapallo · 0 citations
Review Open access 2026

Ordinal Factor Analysis with Robust Estimation and Predictive Models

Ordinal survey indicators are common in behavioral and social science research. Still, they violate continuous normal assumptions when category spacing is unequal, response distributions are skewed, or categories are sparse. This study evaluates a practical workflow for ordinal factor analysis that combines robust confirmatory factor analysis (CFA) estimation with complementary predictive modeling. Simulated five-category ordinal datasets and an empirical Malaysian green consumption dataset were analyzed using WLS, WLSMV, and DWLS estimators based on polychoric correlations. CFA performance was examined through conventional fit indices, parameter recovery, and bootstrap stability of standardized loadings. Predictive models, including random forests, support vector machines, gradient boosting, dense neural networks, convolutional neural networks, and recurrent neural networks, were assessed using matched 5-fold cross-validation. The revised comparison separates measurement validation from prediction, reports estimator stability through resampling, and clarifies the conditions under which robust ordinal estimators and nonlinear predictive models are useful. The study contributes a reproducible benchmark and reporting framework for researchers analyzing Likert-type ordinal data in psychometric and behavioral applications.

A. Osman, Zahayu Binti Md Yusof · 0 citations
Review Open access Aug 2026

Bayesian model comparison for random effects probit models with missing covariates

For model comparison in random effects probit models with incompletely observed covariates, this paper develops a Bayesian data-augmentation workflow in which latent Gaussian responses, random effects, and missing covariate values are updated within a common augmented sampling scheme. Because specifying a fully parametric joint model for mixed continuous and categorical covariates is often unattractive in survey applications, missing covariates are updated by a decision-tree-assisted Bayesian-bootstrap step. Competing models are evaluated conditionally on one common medoid completion using Chib’s method with reduced Gibbs sampling; sensitivity is assessed with respect to the Chib evaluation point, the regression-coefficient prior, and the medoid reference model. The simulation study compares the proposed approach with complete case analysis, multiple imputation by chained equations, missForest single imputation, information criteria, and predictive criteria under MCAR, cross-dependent MAR-type, and self-masked MNAR scenarios. An empirical illustration based on the National Educational Panel Study demonstrates how the method can be used for comparing labor-market models of current employment when competence measures and employment-history covariates are incompletely observed. The results show that missing covariates can materially affect model rankings, and that the proposed workflow provides a transparent evidence-based comparison of nested and non-nested random effects probit specifications under incomplete covariate information.

Michael Bergrab · 0 citations
Open access Sep 2026

Imputing longitudinal PROM data at the item and score level in clinical trials

Missing data in patient-reported outcome measures (PROMs) can affect estimation accuracy and inferential validity, particularly under missing-at-random (MAR) mechanisms. Multiple imputation (MI) is commonly recommended, but the relative performance of item-level and score-level imputation for longitudinal PROM composite scores remains insufficiently understood. We conducted a simulation study based on longitudinal clinical trial PROM data. Predictive mean matching (PMM) and random forest (RF) imputation were evaluated at both the item and score level under empirically calibrated MAR mechanisms. Simulation scenarios varied sample size, overall missingness, and the proportion of unit nonresponse. Performance was assessed using root mean squared error (RMSE), bias magnitude, relative efficiency compared with complete case analysis (CCA), conditional confidence interval coverage, and variance ratio estimates. RMSE decreased with increasing sample size and increased with higher proportions of missingness across all methods. Both PMM and RF generally outperformed CCA in terms of RMSE and relative efficiency, particularly under higher missingness. Differences between item-level and score-level imputation were generally modest, although score-level imputation tended to yield slightly lower RMSE and bias magnitudes under more severe missingness conditions. Conditional coverage remained close to nominal levels across most MI settings. Variance tended to be overestimated in some scenarios, although sensitivity analyses with larger numbers of imputations substantially reduced this effect. PMM showed comparatively stable performance across simulation settings. Multiple imputation methods generally outperformed complete case analysis for handling longitudinal PROM data under MAR. While score-level imputation sometimes showed slightly more favorable performance than item-level imputation when targeting composite score means, the magnitude of these differences was frequently small in practical terms. PMM provided stable performance across a wide range of settings and represents a reasonable default approach for many applied PROM analyses.

Unknown authors · 0 citations
Open access Jul 2026

Academic Performance Forecasting via Data Imputation and Bayesian Neural Networks

This study aims to accurately predict students’ academic performance trajectories for university entrance examinations by proposing a machine learning framework that explicitly accounts for missing data and uncertainty. Mock examination data are characterized by substantial missing values due to heterogeneous participation in exams, as well as inherent randomness caused by variations in test content and examinee conditions. Conventional single-value imputation methods cannot adequately reconstruct the missing values arising from such heterogeneous participation without introducing strong bias, and existing educational prediction models based on deterministic formulations do not account for the inherent randomness and uncertainty in examination scores, thereby limiting the reliability of their forecasts. To address these challenges, we employ GP-VAE and SAITS, state-of-the-art methods for time-series imputation, to reconstruct incomplete mock examination data. Furthermore, we develop a Bayesian Neural Network (BayesNN) to predict future academic performance while explicitly modeling uncertainty. By integrating temporally aware imputation with probabilistic prediction, the proposed framework aims to provide more accurate and reliable performance forecasts than existing approaches. We evaluate the effectiveness of the proposed method through comparative experiments involving various combinations of imputation techniques and prediction models. Experimental results demonstrate that the proposed framework achieves competitive predictive accuracy: the combination of deep imputation methods and BayesNN yields the lowest average estimation error of 15.98 points, compared with 16.75 points for the conventional combination of mean imputation and linear regression. The contribution of this study does not lie in proposing a new deep learning model itself, but rather in systematically comparing combinations of time-series imputation methods and uncertainty-aware prediction models using real-world mock examination sequence data with missing values, thereby providing effective design guidelines for educational data analysis.

Yutaka Yamada, Koshi Watanabe, Keisuke Maeda et al. · 0 citations
Preprint Aug 2026

A Bayesian Longitudinal Model for Imputing Item-Level Missing Data in Trial-Based Economic Evaluations

Trial-based economic evaluations are widely used to assess the cost-effectiveness of healthcare interventions and inform decision-making. Cost and effectiveness outcomes are typically collected using multi-item questionnaires administered at multiple time points, and are often subject to item-level missingness. In principle, imputation (i.e., replacing missing value with estimated or substituted values) should be performed at the item level to fully exploit available information. However, this is rarely implemented in practice due to several statistical challenges, including the longitudinal data structure, cross-item dependence, heterogeneous missingness patterns, and the mixture of skewed cost and count data. In this paper, we develop a Bayesian longitudinal model for imputing item-level missing data in trial-based economic evaluations that accommodates these complexities within a unified framework. The approach combines a transition-model formulation for longitudinal dependence with flexible distributional assumptions and explicit modelling of cross-item relationships, allowing item-level responses of different types to be coherently modelled over time. Motivated by a real-world trial, we demonstrate the flexibility and practical applicability of the proposed approach. We further discuss how the model can be extended to settings where data may be missing not at random.

Xiaoxiao Ling, A. Gabrio, Gianluca Baio · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.