Credit Risk Assessment and Explainability of Machine Learning
Abstract
With the increase in online development of credit business, the customer information that financial institutions can use is also increasing, and the traditional credit scoring model has limitations in fitting high-dimensional characteristics, non-linear correlation, and multi-source heterogeneous data. This paper focuses on traditional statistical methods, machine learning models, and interpretable artificial intelligence in credit risk assessment, focusing on the applicable characteristics of logistic regression, random forest, XGBoost, and neural networks, and discusses the role of interpretation tools such as SHAP and lime in credit audit. The results show that machine learning algorithms have better classification and recognition ability for complex data samples, but the effect of the model will be affected by factors such as sample quality, default caliber, index selection, and business scenarios. At present, there is no single optimal model for all credit tasks. Traditional methods still have irreplaceable value in transparency, stability, and regulatory communication. This paper believes that the credit risk model should choose between forecasting ability, interpretation requirements, misjudgment cost, and compliance conditions, and form a more stable application path through traditional model benchmarks, machine learning assistance, interpretation tools, and manual review.