Skip to content
Open access

Comparative Predictive Performance of Logistic Regression, Naive Bayes, and Support Vector Machine Models in Loan Default Classification among Microfinance Institution Clients in Makueni County, Kenya

2026 · International journal of research and innovation in social science · Vol 10, pp. 8961-8980 · 0 citations

TL;DR

The study concludes that machine-learning models, particularly SVM, significantly improve credit risk prediction in microfinance environments and can support data-driven lending decisions in rural financial institutions.

Abstract

Microfinance institutions (MFIs) play a critical role in improving financial inclusion in Kenya; however, high loan default rates continue to threaten their financial sustainability. This study aimed to develop and compare the performance of Logistic Regression, Naïve Bayes and Support Vector Machine (SVM) models in predicting loan default among clients of MFIs in Makueni County, Kenya. The study adopted a quantitative research design and used secondary data comprising 4,592 borrower records obtained from selected MFIs operating within the county. Data analysis was done using python programming. Borrower socio-economic and financial attributes were extracted from loan records and preprocessed through data cleaning, normalization an. encoding procedures. To ensure robust and unbiased evaluation, Stratified 5-fold cross-validation combined with Grid Search CV hyperparameter tuning was applied across all models. Model performance was assessed using accuracy, precision, recall, F1-score, specificity, and Area Under the Curve (AUC). The results showed that the SVM model achieved the highest predictive performance (accuracy = 0.857, AUC = 0.914), followed by Logistic Regression (accuracy = 0.830, AUC = 0.904), while Naïve Bayes performed least effectively (accuracy = 0.740, AUC = 0.769). The findings demonstrate that SVM provides superior classification ability in capturing complex borrower patterns, while Logistic Regression remains a strong and interpretable baseline model. The study concludes that machine-learning models, particularly SVM, significantly improve credit risk prediction in microfinance environments and can support data-driven lending decisions in rural financial institutions.

Read PDF

Similar papers

Open access Jul 2026

Optimization of Support Vector Machine for Imbalanced Credit Risk Classification

Credit risk classification is an important component of financial decision-making because inaccurate classification may increase payment-default exposure and reduce lending quality. Credit datasets commonly contain numerical variables with different measurement scales and imbalanced distributions between default and non-default clients, which can affect the reliability of machine-learning models. Objective: This study aims to develop and evaluate a Support Vector Machine model for classifying credit card clients into default and non-default categories. The study also examines the influence of numerical standardization, kernel selection, hyperparameter optimization, and balanced class weighting on classification performance. Methodology: A quantitative experimental approach was applied using the Default of Credit Card Clients dataset from the UCI Machine Learning Repository. The dataset consisted of 30,000 observations and 23 predictor variables. Data were divided into training and testing subsets using a stratified 80:20 ratio. Categorical variables were encoded, numerical variables were standardized, and several SVM kernels were evaluated. Hyperparameter selection was conducted using five-fold cross-validation. Model performance was assessed using accuracy, precision, recall, F1-score, specificity, balanced accuracy, and ROC–AUC. Findings: The SVM model trained without standardization failed to identify default clients effectively. Numerical standardization substantially improved classification performance, while the radial basis function kernel produced the strongest validation results. The selected balanced RBF-SVM achieved 77.13% accuracy, 48.54% precision, 56.22% recall, 52.09% F1-score, 83.07% specificity, 69.65% balanced accuracy, and 75.10% ROC–AUC. Balanced class weighting improved default detection but increased false-positive predictions. Implications: The model can support financial institutions as an initial credit-risk screening tool. Its predictions should be combined with document verification, repayment-capacity analysis, and manual assessment rather than being used as the sole basis for credit approval. Originality: This study provides a controlled evaluation of SVM performance by integrating feature standardization, kernel selection, hyperparameter optimization, class-imbalance treatment, and class-sensitive performance metrics. The study demonstrates that credit-risk models should be selected based on balanced default detection rather than overall accuracy alone.

Andre Pratama Adiwijaya · 0 citations
Open access Jul 2026

Non-Performing Loan Prediction for Credit Application Analysis Using Feature Selection and Ensemble Methods

This study aims to identify relevant factors in predicting NPLs and create an NPL prediction model based on these factors and shows hyperparameter tuning was shown to improve recall, thus improving the model's ability to measure how much positive data was successfully predicted by the model.

David Jefri Aruan, Rusdah Rusdah, Ahmad Pudoli · 0 citations
Open access Jul 2026

An Interpretability Analysis of Credit Default Prediction Using Random Forest with SHAP and LIME

This study explores the use of Explainable Artificial intelligence techniques to improve the interpretability of credit default prediction and highlights the practical value of explainable machine learning in developing more understandable, trustworthy, and accountable credit risk assessment systems for real-world financial decision-making.

Muskan, B. Sidhu · 0 citations
Conference Open access Jul 2026

A Comparative Study of Traditional Statistical Models and Machine Learning Algorithms in Credit Risk Assessment

It is suggested that superior ranking performance does not necessarily imply superior decision quality and that effective credit risk modeling requires balancing predictive flexibility with probabilistic reliability and governance stability.

Hanrun Jin · 0 citations
Open access Aug 2026

Toward Reliable Machine Learning Model Selection: A Standardized Multi-Metric Evaluation Framework For Loan Approval

The growing need for accurate, consistent, and reliable loan approval systems The use of machine learning in credit decision-making is increasingly important for financial institutions, but comparative research still often focuses on Accuracy or a limited number of classification metrics, so the trade-off between predictive performance and computational efficiency is not fully described. This study aims to develop a Standardized Multi-Metric Evaluation Framework (MMEF) to support the selection of more objective and reproducible machine learning models in the case of loan approval. The research method uses a standardized experimental pipeline with consistent preprocessing, class balancing using the Synthetic Minority Over-sampling Technique (SMOTE), identical data sharing, model optimization, and multi-metric evaluation. Five algorithms, namely Logistic Regression, Support Vector Machine (SVM), Random Forest, XGBoost, and CatBoost, are compared using a loan approval dataset consisting of 45,000 records and 13 predictor features. The evaluation includes Accuracy, Precision, Recall, F1-score, ROC-AUC, training time, and Overall Score. The results show that XGBoost provides the best overall performance with Accuracy 87.86%, Precision 66.77%, Recall 90.30%, F1-score 76.77%, ROC-AUC 96.27%, and Overall Score 0.827503. CatBoost has the fastest training time of 1.01 seconds, while SVM obtained the highest Recall of 92.40% with a much longer training time. These results indicate that model selection is not sufficient based on a single metric. MMEF provides a more systematic evaluation basis to identify models that have a balance of performance and efficiency in loan approval experiments.

Trihartono Agus, Agus Ilyas Ilyas, S. Sattriedi et al. · 0 citations
Open access Jul 2026

Optimizing Credit Risk Assessment in Ghanaian Micro-Lending Institutions: A Comparative Analysis of Random Forest, Extra Tree Classifier, and Ensemble Machine Learning Models

It is concluded that integrating ML models can substantially improve the accuracy, consistency, and reliability of credit risk evaluations, thereby reducing default rates and supporting financial inclusion in Ghana's microfinance sector.

P. Addo, Samuel Kofi Akpatsa, Emmanuel Mensah et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.