Skip to content
Review Open access

A Comprehensive Review on Missing Data Imputation Techniques

2026 · ITEGAM- Journal of Engineering and Technology for Industrial Applications (ITEGAM-JETIA) · Vol 12, pp. 1525-1532 · 0 citations

TL;DR

The results indicate that machine learning algorithms, particularly missForest and KNN, consistently provide high accuracy by modeling complex and nonlinear relationships and recommend advanced machine learning algorithms to ensure unbiased inference, while cautioning against simple imputation techniques that may introduce bias.

Abstract

Missing data is a fundamental challenge in scientific research and often leads to biased results and weakened statistical inference. To address this, imputation methods are essential to preserve the original information. This study compares a wide range of imputation techniques categorized into two main types: traditional statistical methods (mean, median, mode, last observation carried forward, regression, multiple imputation), and machine learning and deep learning methods including algorithms such as (k-nearest neighbors, random forests, support vector machines, neural networks, generative adversarial networks). The results indicate that machine learning algorithms, particularly missForest and KNN, consistently provide high accuracy by modeling complex and nonlinear relationships. Furthermore, GAN-based methods such as GAIN, GIMIN, and MisGIMIN represent significant advancements, especially for high-dimensional data with missing rates exceeding 80%. Specifically, the GIMIN algorithm demonstrated superior performance in root mean square error (RMSE) accuracy at high levels of missingness, while MisGAN achieved the lowest FID error. The study emphasizes selecting methods based on the characteristics of the dataset and recommends advanced machine learning algorithms to ensure unbiased inference, while cautioning against simple imputation techniques that may introduce bias.

Read PDF

Similar papers

Review Open access Jul 2026

A Systematic Literature Review of Missing Data Imputation Techniques in Tabular Machine Learning Datasets

Losses of data are a widespread issue of the real-world tabular data, utilized in machine learning (ML). Missing values may dramatically hamper the quality of the model, be biased, and result in incorrect inferences unless addressed correctly. This is a systematic literature review (SLR) that explores and syntheses 52 research articles published 2020-2026 in high-impact peer review journals. The review is done under the guidelines of PRISMA (Preferred Reporting Items to Systematic Reviews and Meta-Analyses). Methods of imputation can be divided into 5 broad categories: statistical and conventional imputation methods (mean, median, mode, and Last Observation Carried Forward), machine learning-based methods (k-Nearest Neighbors, Random Forest, Decision Trees and Support Vector Machines), multiple imputation methods (including MICE, missForest and missRanger), deep learning-based methods (including Autoencoders, Vari There is a systematic comparison between methods based on type of dataset, missing data mechanism (MCAR, MAR, MNAR), evaluation measures (RMSE, MAE, accuracy, AUC), computational complexity and scalability. Using our results, it appears that, up to low missingness rates, conventional approaches are equally competitive, but that deep generative models (with GAN-based models or diffusion-based models being two different approaches to the same task) are matched when applied to high-dimensional and heterogeneous tabular data. However, there is no one particular approach that prevails in all situations. This review finds the overall gaps in research, such as the absence of standardized benchmarks, the relative dearth of interest in MNAR mechanisms, and the lack of research on imputation in federated learning. The results give practical advice to practitioners and researchers to use the right imputation techniques when using tabular ML tasks.

Nabeel Ali Khan, Munir Ahmad, Shamila Ghafoor et al. · 0 citations
Open access Jul 2026

Not All Missing Data are Equal: Choosing the Right Imputation Method for Binary Datasets

Missing binary predictors are common in reliability, quality control, and industrial decision systems, yet imputation methods are often chosen by convenience rather than evidence. We conduct a Monte Carlo study comparing mode substitution, sequential hot‐deck, missForest, MICE, and KNN with three neighbourhood sizes under MCAR, MAR, and MNAR missingness, across missingness rates from 5% to 50% and two predictor‐dependence structures. Performance is evaluated on three targets: exact recovery of missing binary cells, recovery of logistic‐regression coefficients, and downstream classification using logistic regression, naive Bayes, support vector machines, and random forests. The results reveal a clear trade‐off. KNN is strongest for exact cell recovery under MCAR and MAR, whereas missForest performs best under MNAR. MICE is the most reliable choice for downstream predictive performance across learners and missingness mechanisms. By contrast, mode imputation and sequential hot‐deck achieve the best coefficient recovery. The main implication is operational: in binary‐data environments, imputation should be chosen to match the analytical objective–reconstruction, inference, or prediction–because no single method dominates all targets simultaneously.

Manuel Delfino, Fabio Rapallo · 0 citations
Open access Aug 2026

Machine Learning-Based Imputation for Breast Cancer Prediction: Evaluating Performance Under Complex Missing Data Mechanisms

This study systematically compared statistical and machine learning-based imputation methods using two publicly available breast cancer datasets representing complementary clinical settings to highlight the importance of considering dataset characteristics, missing-data mechanisms, and the intended analytical objective when selecting imputation methods.

Nyatuga Gideon Nyakundi, John Ndiritu, Ivivi J. Mwaniki et al. · 0 citations
Jul 2026

Enhancing Medical Data Imputation Using a Denoising Autoencoder with Missing-Neighborhood Perturbation.

In real-world clinical settings, the diverse types and unknown causes of missing data in tabular medical datasets pose significant challenges for accurate imputation. In particular, non-random missingness-where missing values are related to unobserved variables-limits the effectiveness of many existing imputation models. To address these challenges, we propose a robust and generalized imputation method: Multiple Imputation based on Neighborhood Perturbation Denoising Autoencoder (MI_NPDAE). MI_NPDAE identifies optimal donor records by leveraging neighborhood information, which is used as input to the autoencoder. The model reconstructs perturbed inputs to learn robust feature representations around missing regions, while the introduction of additive noise exposes the model to a variety of missingness patterns, enhancing its adaptability. We evaluate MI_NPDAE on two datasets: a publicly available Breast dataset and a lung cancer nutrition dataset from the Chinese Anti-Cancer Society. Experimental results demonstrate that MI_NPDAE consistently outperforms baseline methods across various missing mechanisms and ratios, maintaining lower imputation errors. Moreover, the imputed data significantly improves performance in downstream predictive tasks, highlighting the practical value of our approach in clinical data analysis.

Huamei Qi, Chen Cao, Wenhui Yang et al. · 0 citations
Open access Aug 2026

A GAN-Enhanced and Cluster-Aware Data Preprocessing Framework for Robust Predictive Healthcare Analytics

Missing values and significant class imbalance are common characteristics of healthcare datasets which significantly impair predictive model performance and reduce their dependability in clinical decision-making. Creation of reliable and broadly applicable healthcare prediction system depends on addressing this issues As to improve overall data quality, this study suggests an integrated data preparation system that integrates cluster-aware oversampling methods with Generative Adversarial Imputation Networks (GAIN). By using adversarial training to understand intricate underlying data distributions GAIN model are used to estimate missing values while maintaining significant statistical correlations between variables. Simultaneously, hybrid SMOTE-ENN method are used to remove ambiguous and noisy data and efficiently handle class imbalance. Real-world diabetic readmission dataset are used to assess suggested methodology, and show notable gains in data completeness distribution preservation, and prediction performance. Significant improvements in accuracy, recall, and F1-score are revealed by experimental data, suggesting improved capacity to detect high-risk individuals. As compared to traditional methods incorporation of sophisticated preprocessing technique enhances model resilience and generalisation. This results highlight significance of integrating class balancing technique and intelligent imputation into single framework. Overall, study emphasises how important sophisticated preprocessing are to enhancing clinical applicability, robustness and dependability of predictive healthcare analytics system.

Aminu Usman Jibril, A. S. Kumar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.