Payroll-anchored retail borrowers—individuals whose monthly remuneration is routed into an account at the lending institution through a salary-project arrangement—constitute the volume backbone of unsecured consumer lending in Kazakhstan, generating the largest origination flow, the lowest realized default rate, and the majority of the systemic regulatory and capital sensitivities of second-tier banks. Payroll anchoring also changes the lender’s information set, which motivates a study of how that advantage translates into model performance and borrower outcomes. We design and internally validate an explainable hybrid artificial-intelligence framework stratified by client tenure into two production models: a Weight-of-Evidence (WOE) logistic-regression scorecard for new salary-project applicants, and a hybrid scorecard for repeat applicants, in which a stacked ensemble of LightGBM, CatBoost and a multi-head self-attention neural network contributes a single WOE-encoded predictor to a second-stage L2-regularized logistic regression. The hybrid recovers a substantial share of the ensemble’s discriminatory lift while preserving an auditable, monotone scorecard at the point of decision, and isotonic recalibration restores the predicted probabilities of default to the empirical bad-rate scale required for IFRS 9 expected-credit-loss accrual and risk-based pricing. We report discrimination, calibration and stability evidence under a strict anti-leakage protocol and set out the structural preconditions under which the architecture transfers to other emerging-market payroll-anchored portfolios. We are explicit about scope: a true out-of-time validation and a full group-conditional fairness audit are identified as required next steps rather than claimed here. The contribution is a reproducible, interpretable scoring design that exploits payroll visibility while retaining full coefficient interpretability inside the production decision engine.
Real-time retail credit scoring is a data-intensive cognitive computing task. Each decision must fuse heterogeneous signals, execute a non-linear model, return a calibrated probability of default (PD), and emit a regulator-compliant local explanation within milliseconds. We address the most demanding segment of unsecured lending in Kazakhstan—Salary-Project-Independent (SPI) borrowers, whose principal income stream is not observable by the lender—and frame scoring as a constrained optimisation problem where we maximise discrimination subject to interpretability, latency, and calibration constraints. We propose a tenure-stratified hybrid framework that couples (i) an online weight-of-evidence logistic regression (WOE-LR) scorecard with (ii) an offline self-attention stacked ensemble (LightGBM, CatBoost, and a tabular self-attention network) whose calibrated PD is quantile-binned, WOE-encoded, and re-injected into the online scorecard as a single auditable predictor. On 551,962 production contracts that originated in 2022–2024, the repeat-client hybrid attains an area under the receiver operating characteristic curve (AUROC) of 0.826, a Gini coefficient of 0.65, and a Kolmogorov–Smirnov (KS) statistic of 0.495, preserving roughly half of the offline ensemble’s lift over the linear baseline (AUROC 0.79→0.897) while retaining a fully auditable twelve-coefficient scorecard in production. The new-client scorecard attains an AUROC of 0.741. Non-parametric isotonic recalibration reduces the expected calibration error from 0.27 to below 0.01 and raises the Hosmer–Lemeshow p-value above 0.99 without altering discrimination. The framework complies with the model risk standards of the Agency of the Republic of Kazakhstan for Regulation and Development of the Financial Market and is delivered as a Spark/MLOps reference architecture, illustrating how big data engineering, attention-based representation learning, and post hoc explanations can be co-designed for a high-stakes, high-throughput, regulated AI application.
Gulnaz Zakariya, A. Moldagulova, Nor’ashikin Ali· Big Data and Cognitive Compu...· 0 citations
SME credit files arrive incomplete, and the gaps leave room for anchoring, confirmation bias and loss aversion. We compare human, algorithmic and hybrid credit decisions using a benchmark credit dataset alongside a vignette experiment with credit analysts working in Morocco's Souss-Massa region. The modelling arm pairs L2-regularised logistic regression with gradient-boosted trees, adding stratified validation, calibration analysis, SHAP and LIME. In the human arm, matched cases vary the requested amount while everything else is held constant. Analysts were least stable on borderline files, and their decisions moved with the anchor. The boosted model held steadier but leaned harder on indicators that track how thick a file is. AI-first assistance improved consistency and deepened deference to the model; human-first assistance preserved contextual overrides; explanation-gating struck the best balance, though only where SHAP and LIME agreed. We assess distribution through demographic-parity difference, disparate-impact ratio, equal-opportunity difference and false-positive-rate difference. What the results support is a governed hybrid: weak explanations withheld, overrides auditable, human review genuinely available. A regional sample and benchmark data bound how far any of these travels.
Hassan Ennaqui, Mohamed El Bourki, Abdellah Bakrim et al.· EPJ Web of Conferences· 0 citations
Probability-of-default (PD) estimation under the Basel and IFRS 9 frameworks requires probabilities that are both discriminative and well-calibrated. Large language models applied to credit assessment via prompting or supervised fine-tuning (SFT) yield poorly calibrated probabilities, while reinforcement learning with binary correctness rewards is structurally unsuitable for probability prediction under extreme class imbalance. We propose CreditR1, a three-stage framework: an SFT cold start on evidence-filtered reasoning chains; Group Relative Policy Optimization, guided by a composite verifiable reward combining Brier-score calibration—a strictly proper scoring rule—pairwise ranking, evidence anchoring, and format compliance; and an anti-contamination evaluation protocol. On Chinese A-share corporate credit data, CreditR1 matches gradient-boosted baselines in discrimination (AUC: 0.883±0.004 vs. 0.891 for XGBoost) while reducing expected calibration error by 24.2% versus isotonic-calibrated XGBoost (ECE: 0.047 vs. 0.062) and by 47.2% versus uncalibrated XGBoost (0.089). Because the test set contains only 119 default events, all comparisons carry firm-level bootstrap confidence intervals; the calibration advantage remains significant against every baseline after Holm–Bonferroni correction, including Platt, beta, and Bayesian-binning recalibrations. Ablations confirm each reward component is necessary. CreditR1 delivers calibrated PDs with evidence-grounded reasoning that supports internal model validation and human review; transferability beyond the Chinese A-share market remains an open empirical question.
The cooperative banking sector in Nepal constitutes a foundational pillar of financial inclusion, serving approximately 7.4 million members across 31,450 primary cooperatives and disbursing loans exceeding NPR 453 billion. Yet this sector is hemorrhaging credibility. The National Cooperative Bank Limited reported a non-performing loan ratio of 33.01 percent as of mid-July 2025, with its capital adequacy ratio collapsing to 0.82 percent — a figure that would trigger immediate regulatory intervention in any conventional banking jurisdiction. Against this backdrop, the question is no longer whether cooperative banks in Nepal need better risk assessment tools, but whether they can afford to continue without them.
This paper develops a comprehensive theoretical framework for integrating explainable machine learning into credit risk management systems of Nepalese cooperative banks. We derive the complete mathematical architecture of ensemble gradient boosting — specifically XGBoost and LightGBM — alongside post-hoc interpretability mechanisms grounded in cooperative game theory, namely SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations). The framework incorporates information-theoretic feature selection, SMOTE-based class imbalance correction, fairness constraints calibrated to Nepal’s socio-economic heterogeneity, and a rigorous convergence analysis of the boosting iteration. We prove consistency of the mutual-information feature selector, derive the bias-variance decomposition for gradient-boosted ensembles, establish generalization bounds via Rademacher complexity, and verify the four Shapley axioms for the TreeSHAP algorithm. Drawing upon Nepal Rastra Bank’s Financial Stability Report FY 2024/25, the National Cooperative Federation of Nepal’s sectoral statistics, and the Government of Nepal’s Economic Survey 2025/26, we situate the technical apparatus within Nepal’s regulatory architecture — the Cooperative Act 2017, the Financial Sector Development Strategy 2022–2026, and the National Financial Inclusion Roadmap. The paper argues that predictive accuracy and regulatory transparency are not competing objectives but complementary necessities for institutional survival in Nepal’s cooperative sector.
S. K. Sahani, Tsair-Fwu Lee, Digvijay Pandey et al.· Journal of Intelligent Decis...· 0 citations
Credit risk assessment remains a foundational function of commercial banking, yet the Indian banking sector's persistent non-performing asset (NPA) burden, heterogeneous borrower base, and expanding priority-sector and microfinance lending create distinct challenges that generic, globally trained credit scoring models do not adequately address. This paper proposes an Explainable Machine Learning (XML) framework for credit risk assessment that combines an ensemble classifier, integrating XGBoost, Random Forest, and LightGBM, with an integrated SHAP-and-LIME explainability layer, and evaluates the framework using a large-scale retail and priority-sector loan dataset drawn from public sector, private sector, regional rural, and small finance bank segments operating in India. The framework incorporates SMOTE-ENN based class-imbalance handling to address the low base default rate typical of retail lending portfolios, and an explainability-guided feature refinement step that uses SHAP attributions to iteratively prune low-value features and retain regulator-interpretable risk factors. The proposed model was evaluated on a dataset of 42,860 loan accounts and benchmarked against five baseline models. Results show that the proposed explainable ensemble achieves 93.2% accuracy, 91.0% precision, 89.4% recall, and a 90.2% F1-score, exceeding the strongest baseline (LightGBM) by 5.6 percentage points in F1-score, with an area under the ROC curve (AUC) of 0.967. Segment-wise analysis reveals materially higher default risk concentration in regional rural banks and small finance institutions relative to public and private sector banks, with debt-to-income ratio, credit bureau (CIBIL) score, and repayment delinquency history emerging as the most influential predictors across segments.
Amit Kumar Agrawal, Vaibhav Gandhi· International journal of com...· 0 citations
Artificial intelligence-integrated ‘buy now, pay later’ (BNPL) platforms are diffusing rapidly across the Middle East and North Africa (MENA), raising concerns about consumer financial vulnerability. Drawing on choice architecture, payment decoupling, and financial literacy literatures, this study examines how three platform-level features—algorithmic nudging, AI personalization intensity, and perceived ease of credit—are associated with impulsive buying tendency and downstream financial outcomes, and whether BNPL-specific financial literacy attenuates these associations. A multi-method design combined cross-sectional partial least squares structural equation modeling (N = 1247 active BNPL users in seven MENA countries) with a six-month longitudinal follow-up (N = 847, 68% retention). Algorithmic nudging was positively associated with impulsive buying tendency, which in turn was associated with elevated financial stress and longitudinal debt accumulation. The ‘loyalty trap’—a paradoxical state in which financially stressed consumers maintain high platform loyalty—is provisionally documented via piecewise longitudinal trajectories. We emphasize that this pattern is consistent with but not causally established by the present design, and we outline specific experimental and quasi-experimental research designs needed for causal identification. BNPL-specific financial literacy moderated the associations between algorithmic nudging, impulsive buying, and adverse financial outcomes, with the highest-literacy quartile exhibiting substantially attenuated debt trajectories. We discuss boundary conditions, alternative explanations, and the limits of causal inference in non-experimental panel data. Findings inform evolving BNPL regulatory frameworks in MENA, with particular relevance to nudge-transparency disclosures, contractual cooling-off periods, and credit-bureau reporting standards.
Osama Wagdi, Walid Abouzeid, Heba Farid et al.· Journal of Theoretical and A...· 0 citations