An Explainable Machine Learning Framework for Early Detection of Chronic Kidney Disease Using Optimized SVM
Abstract
The early detection of Chronic Kidney Disease (CKD) is critical in order to commence the right treatment and minimize the chances of the disease from progressing. This paper outlines an interpretable machine learning algorithm for chronic kidney disease classification based on a fine-tuned Support Vector Machine (SVM). In this research, we use a clinical dataset that comprises 400 patient records and medical attributes. Data preprocessing involved the management of missing values, normalization of numerical features, and encoding of categorical features. The issue of class imbalance was dealt with by applying Synthetic Minority Oversampling Technique (SMOTE). For determining the optimal hyperparameters for the SVM, we followed an exhaustive grid-search approach using GridSearchCV. Model performance was evaluated based on accuracy, precision, recall, F1 score, and ROC-AUC. Experiment results have shown that the fine-tuned SVM has reached 98.75% accuracy and 0.9993 ROC-AUC. Moreover, a mean accuracy of 99.75% in ten-fold cross-validation suggests that the model is consistent in its predictions on various subsets of data. In order to increase the interpretability of our model, SHapley Additive exPlanations (SHAP) were included into the framework to evaluate the importance of the specific clinical features in the model's predictions.