A GAN-Enhanced and Cluster-Aware Data Preprocessing Framework for Robust Predictive Healthcare Analytics
Abstract
Missing values and significant class imbalance are common characteristics of healthcare datasets which significantly impair predictive model performance and reduce their dependability in clinical decision-making. Creation of reliable and broadly applicable healthcare prediction system depends on addressing this issues As to improve overall data quality, this study suggests an integrated data preparation system that integrates cluster-aware oversampling methods with Generative Adversarial Imputation Networks (GAIN). By using adversarial training to understand intricate underlying data distributions GAIN model are used to estimate missing values while maintaining significant statistical correlations between variables. Simultaneously, hybrid SMOTE-ENN method are used to remove ambiguous and noisy data and efficiently handle class imbalance. Real-world diabetic readmission dataset are used to assess suggested methodology, and show notable gains in data completeness distribution preservation, and prediction performance. Significant improvements in accuracy, recall, and F1-score are revealed by experimental data, suggesting improved capacity to detect high-risk individuals. As compared to traditional methods incorporation of sophisticated preprocessing technique enhances model resilience and generalisation. This results highlight significance of integrating class balancing technique and intelligent imputation into single framework. Overall, study emphasises how important sophisticated preprocessing are to enhancing clinical applicability, robustness and dependability of predictive healthcare analytics system.