Skip to content
Open access

Non-Performing Loan Prediction for Credit Application Analysis Using Feature Selection and Ensemble Methods

Jul 2026 · IDEALIS : InDonEsiA journaL Information System · Vol 9, pp. 430-439 · 0 citations · 20 references

TL;DR

This study aims to identify relevant factors in predicting NPLs and create an NPL prediction model based on these factors and shows hyperparameter tuning was shown to improve recall, thus improving the model's ability to measure how much positive data was successfully predicted by the model.

Abstract

Non-Performing Loans (NPL) are a fundamental indicator of a financial institution's asset health, reflecting loans that fail to meet interest or principal payment obligations as agreed. A high NPL ratio negatively impacts a bank's financial performance, such as decreased profitability as measured by Return on Assets (ROA) and decreased liquidity. Bank Indonesia sets an NPL tolerance limit of 5% of total credit provided by banking financial institutions. Therefore, a predictive model is needed that can detect the possibility of customers experiencing NPLs early. This study aims to identify relevant factors in predicting NPLs and create an NPL prediction model based on these factors. The contribution of this study lies in combining the results of three feature selection techniques: Chi-Square, Mutual Information, and Random Forest feature importance, using the average score eliminated by the Recursive Feature Elimination technique. Several ensemble algorithms, namely Random Forest, XGBoost, Gradient Boosting, and LightGBM, were explored to produce the best-performing model. Then, hyperparameter tuning was performed on the best model. The Random Forest model produced the best performance, with 92.17% accuracy, 78.1% precision, 98.1% recall, and 95.5% AUC. Hyperparameter tuning was shown to improve recall, thus improving the model's ability to measure how much positive data (Current class) was successfully predicted by the model. The results of this study can assist management in making credit decisions. Thus, it is hoped that it can help reduce the number of NPL cases.

Read PDF

Similar papers

Open access Jul 2026

Application of Binary Logistic Regression to Identify Determinants of Non-Performing Loans

This study aimed to identify the factors associated with non-performing loans (NPLs) in banks and develop a borrower-level model to estimate the probability of loan default. Unlike previous studies that mainly focus on bank-level financial indicators or macroeconomic factors, this study utilizes borrower characteristics and loan information obtained from credit application data at a commercial bank in Palembang, Indonesia. The study used data from 100 borrowers, with predictor variables including age, number of family dependents, total household income, occupation, educational attainment, loan amount, loan term, and monthly installment amount. Binary logistic regression with backward elimination was applied to identify significant predictors of NPLs and to estimate the probability of loan default. The results showed that occupation, loan amount, and loan term significantly influenced the occurrence of non-performing loans. The final model achieved a classification accuracy of 81% and an area under the receiver operating characteristic curve (AUC) of 0.858, indicating good predictive performance. The obtained binary logistic regression model can be used to estimate the probability of non-performing loans and assist banks in identifying potential credit risks during the credit evaluation process.

Putri Sari Nilam Cayo, Ngudiantoro, Irmeilyana · 0 citations
Open access Aug 2026

Construction of SME Loan Default Risk Prediction System for Commercial Banks Using XGBoost Model

Small and medium-sized enterprises play an important role in promoting employment and innovation, but their small scale, opaque financial information, and unstable operating conditions increase credit risk for commercial banks. Accurate prediction of SME loan default risk can reduce credit losses, optimize credit-resource allocation, and support sustainable SME development. Taking SME loan data from commercial banks as the research object, this paper constructs a loan default risk prediction system based on the XGBoost model. Key indicators affecting loan default are first identified through literature research and expert interviews, including enterprise financial indicators, non-financial indicators, and macroeconomic indicators. The collected data are then cleaned, missing values are processed, and feature engineering is conducted. The XGBoost model is constructed and optimized through grid search and cross-validation, and compared with logistic regression, random forest, and LightGBM models. Evaluation results based on confusion matrix, ROC curve, and AUC show that the XGBoost model achieves an AUC of 0.89 and maintains an AUC of 0.88 on the independent test set, indicating strong predictive accuracy, stability, and non-overfitting performance.

F. Yan · 0 citations
Open access Aug 2026

Predicting Credit Risk with ESG Factors Using XGBoost and Structural Learning in Vague Environments (SLAVE) in Commercial Banks

Predicting credit risk is vital for banks as it safeguards financial stability, minimizes default losses, optimizes capital, and ensures regulatory compliance. This study aims to predict credit risk (High/Low) in commercial banks by integrating machine learning with traditional econometric approaches. The Structural Learning in Vague Environments (SLAVE) fuzzy rule-based model handles ambiguity in financial decisions, while the eXtreme Gradient Boosting (XGBoost) uncovers non-linear patterns among predictors. Input variables—profitability, liquidity risk, ESG (environmental, social, and governance) score, and monetary freedom—were selected via multicollinearity tests and three panel regression models, including ordinary least squares (OLS), fixed effects, and random effects models. The empirical investigation uses a panel dataset of forty commercial banks across seven Middle Eastern countries from 2014 to 2023, yielding 400 observations. Regression results reveal that profitability and ESG score significantly reduce credit risk. Liquidity risk and monetary freedom increase credit risk. XGBoost combined with the SHapley Additive exPlanations (SHAP)-based interpretation identifies ESG Score as the most influential predictor. The SLAVE model was evaluated using three data splits: 70/30, 80/20, and 90/10. The 80/20 split achieved the highest accuracy, with superior performance in identifying low-risk banks. Stronger ESG performance and stable monetary environments contribute to fostering sustainable banking and reducing credit risk, making these indicators valuable for risk management frameworks in the Middle Eastern banking sector.

Jamil J. Jaber, A. A. Alkhawaldeh, Qusay Ayman Sulayman Mazahreh et al. · 0 citations
Open access Jul 2026

An Interpretability Analysis of Credit Default Prediction Using Random Forest with SHAP and LIME

This study explores the use of Explainable Artificial intelligence techniques to improve the interpretability of credit default prediction and highlights the practical value of explainable machine learning in developing more understandable, trustworthy, and accountable credit risk assessment systems for real-world financial decision-making.

Muskan, B. Sidhu · 0 citations
2026

Credit Worthiness Prediction Model Using Artificial Neural Network

The study focused on developing a creditworthiness prediction model utilizing artificial neural network. Credit risk evaluation has a relevant role for financial institutions, as lending could result in real and immediate losses. In particular, default prediction was one of the most challenging activities in managing credit risk. The objective was to enhance the accuracy and reliability of credit risk assessments by leveraging the computational power and learning capabilities of artificial neural networks. The parameters of the dataset include the customer's place of work, loan history, monthly salary, loan amount, transaction history, and credit history, all stored in the trained database. When a customer comes to apply for a loan, the system checks if the user is qualified based on these parameters and then approve or disapprove the loan accordingly. Rigorous testing and validation were conducted to ensure the model's robustness and generalizability. The results demonstrated that the neural network-based model significantly outperformed traditional statistical methods, providing more precise predictions of creditworthiness. An Object-Oriented Analysis and Design Methodology (OOADM) approach, which incorporated Unified Modeling Language for analysis and design, was used. The development stage was completed using a set of software tools, including Python and the MySQL database system.

E. C., Uzo Blessing Chimezie, Ukekwe Emmanuel C · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.