Optimizing Credit Risk Assessment in Ghanaian Micro-Lending Institutions: A Comparative Analysis of Random Forest, Extra Tree Classifier, and Ensemble Machine Learning Models
Jul 2026· Journal of Applied Social Sciences and Entrepreneurship Education· Vol 1, pp. 25-44· 0 citations
TL;DR
It is concluded that integrating ML models can substantially improve the accuracy, consistency, and reliability of credit risk evaluations, thereby reducing default rates and supporting financial inclusion in Ghana's microfinance sector.
Abstract
Credit risk assessment is pivotal to the sustainability of micro-lending institutions, particularly in emerging economies such as Ghana, where conventional evaluation methods remain predominantly manual and subjective. Traditional approaches, which rely on face-to-face interviews, personal judgments, and simple background checks, are vulnerable to human biases, inconsistencies, and inefficiencies that contribute to elevated default rates and broader financial instability. This study investigates the application of machine learning (ML) techniques, specifically Random Forest (RF), Extra Tree Classifier (ETC), and a probability-averaged Ensemble Classifier, to enhance credit risk assessment in Ghanaian micro-lending institutions. Using a quantitative experimental research design, the study analysed 32,581 loan records drawn from Tepa Man Microfinance Institution. Data preprocessing included missing-value imputation, one-hot encoding, and class balancing via random oversampling, applied exclusively to the training set. Model performance was evaluated through 10-fold stratified cross-validation using accuracy, precision, recall, F1-score, AUC-ROC, Cohen's Kappa, and Matthews Correlation Coefficient (MCC). Hyperparameters were set to scikit-learn defaults (n_estimators = 100, random_state = 42) to ensure reproducibility. The Random Forest and Extra Tree Classifiers each achieved a mean accuracy of 99.33% and an AUC-ROC of 0.9997, results that are consistent with the high-quality, real-world dataset and are critically interpreted in the context of potential overfitting risks. Feature importance analysis identified the loan-to-income ratio and interest rate as the dominant predictors of default. The Ensemble Method, which averages class probabilities across both base models, achieved 84.25% accuracy and an AUC of 0.9231, demonstrating stronger generalization than the individual classifiers. The study concludes that integrating ML models can substantially improve the accuracy, consistency, and reliability of credit risk evaluations, thereby reducing default rates and supporting financial inclusion in Ghana's microfinance sector.
An Explainable Machine Learning (XML) framework for credit risk assessment that combines an ensemble classifier, integrating XGBoost, Random Forest, and LightGBM, with an integrated SHAP-and-LIME explainability layer is proposed and evaluated using a large-scale retail and priority-sector loan dataset drawn from public sector, private sector, regional rural, and small finance bank segments operating in India.
A. Agrawal, Vaibhav C. Gandhi· International journal of com...· 0 citations
It is suggested that superior ranking performance does not necessarily imply superior decision quality and that effective credit risk modeling requires balancing predictive flexibility with probabilistic reliability and governance stability.
A systematic literature review of ML applications in credit risk assessment (CRA), covering publications from January 2016 to May 2026, synthesise prevailing methodologies into a unified end-to-end credit risk modelling framework spanning data preprocessing, feature engineering, model training, evaluation, and operational deployment.
Bolun Zhang, Jun Luo, Ruobing Wu et al.· Journal of Risk and Financia...· 0 citations
This study explores the use of Explainable Artificial intelligence techniques to improve the interpretability of credit default prediction and highlights the practical value of explainable machine learning in developing more understandable, trustworthy, and accountable credit risk assessment systems for real-world financial decision-making.
Muskan, B. Sidhu· International Journal of Com...· 0 citations
Banking is increasingly shaped by expanding data volumes, more complex borrower behaviour, and stricter credit risk management requirements. Under such conditions, scoring models are becoming especially relevant as instruments for the formalised assessment of creditworthiness, combining analytical accuracy, speed of decision-making, and the possibility of integration into the bank’s risk management system. The study compares traditional and modern scoring models in bank credit risk management and proposes an approach to their practical use in Ukrainian banking. Its focus is on scoring models as instruments for credit risk assessment. The study combines comparative analysis, matrix modelling, simulation, statistical modelling, and machine learning methods. Given limited access to primary banking information and confidentiality requirements, the empirical analysis was conducted on a synthesised demonstration dataset designed to reflect the structure of a real retail credit portfolio. For the analysis, a sample of 1,000 observations with a default share of 22.0% was constructed, and logistic regression, discriminant analysis, Random Forest, XGBoost, and a hybrid logit + ML re-ranking model were used for comparison. The results showed that XGBoost provided the highest predictive accuracy, with an AUC-ROC of 0.861, Gini of 0.722, Recall of 0.781, and Brier score of 0.141, whereas logistic regression demonstrated an AUC-ROC of 0.781 and retained advantages in terms of interpretability and suitability for validation. The hybrid model achieved an AUC-ROC of 0.848, Gini of 0.696, Recall of 0.773, and Brier score of 0.144, thus ensuring the best balance between accuracy, explainability, calibration, and practical applicability. Practically, the study offers an adaptive approach to selecting scoring models and a matrix for evaluating them under Ukrainian banking conditions, taking into account the requirements of the regulatory environment, data quality, and the instability of the operating conditions of Ukrainian banks.
By empirically proving that high-performance algorithms can be mathematically blind to demographic biases, this framework directly advances SDG 10 (Reduced Inequalities) and provides the accountable, feature-level justifications required for secure and sustainable financial inclusion (SDG 8).
Htet Nge Nge Ko, Aung Htoo Khine, Shadab Kalhoro et al.· Journal of Risk and Financia...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.