An Extensive Survey of Data Mining Methods and their New Applications in Loan Default Detection: Challenges and Future Research Directions
Abstract
Data mining is an effective analytical tool that has become an emerging technique to discover meaningful patterns, relationships, and prediction from large-scale data sets. Data mining is being used in various fields including healthcare, education, business, cybersecurity, agriculture, and finance to enhance decision-making processes in the face of the fast-growing digital data. Different machine learning and statistical methods such as classification, clustering, association rule mining as well as predictive modelling have been extensively utilized in knowledge discovery and automated decision support. Financial risk assessment is one of the several applications that have been in the spotlight because of the availability of customer transaction and credit related data. Prediction of default is a very critical issue for financial institutions as improper credit appraisal can lead to financial risk and economic loss. While multiple data mining techniques have been attempted for credit risk analysis, there are still some issues to be studied, namely: imbalance of the data sets, interpretability of the model, new borrower behavior, and limitations for using the models in real scenarios. This paper provides full overview of the current applications of data mining and discusses how they have helped in intelligent decision-making systems. Moreover, the study suggests that loan default detection is a potential research field where advanced data mining techniques can be developed for an accurate, reliable and explainable loan default prediction model. The review also identifies potential avenues for further research and development in the field of loan default prediction, which include the incorporation of hybrid algorithms, deep learning techniques, explainable artificial intelligence, and real-time financial data analysis.