Diabetes forecasting by analyzing electronic health record data
Abstract
In recent years, the prevalence of diabetes has surged, presenting a significant challenge to healthcare systems worldwide. To address this issue, innovative approaches are imperative to enhance prediction and management capabilities. Consequently, this study underscores the potential of utilizing Electronic Health Records (EHRs) for diabetes prediction through data mining methods. The primary objectives entail forecasting diabetes based on EHRs using Data Mining and Machine Learning techniques, followed by the analysis of crucial features associated with diabetes risk. To achieve these aims, the research leveraged EHR data sourced from The Korean Health Records of the Korean Genome and Epidemiology (KoGES), comprising five data files spanning five different time points. Machine Learning models were then employed for diabetes forecasting. These files were subdivided into four scenarios, each representing forecasting with varying numbers of time points (2, 3, 4, and 5), aimed at assessing diabetes prediction across multiple scenarios and comparing model performances. Then, Machine Learning models were applied to the processed data for forecasting across these scenarios. The final phase involved feature analysis utilizing the SHAP method. The results indicate that the models exhibited optimal performance in the scenario with five time points, yielding AUC-PR and Weighted F1 scores of 0.94 and 0.97, respectively, for the Random Forest model. Additionally, the model demonstrated commendable Precision and Recall metrics for each class, especially in diabetes class, with 0.86 and 0.75 respectively. The analysis of important features revealed associations with future diabetes risk. These outcomes suggest that the models were effectively constructed and trained, exhibiting relatively high predictive precision. The analysis of important features revealed associations with future diabetes risk. These outcomes suggest that the models were effectively constructed and trained, exhibiting relatively high predictive precision. However, certain challenges remain, including suboptimal model performance and the need to improve SHAP's interpretability concerning features of medium and small importance.