Skip to content
Preprint

$\texttt{findr}$: Transparent and Fair Credit Risk Decisions through Semi-Structured Regressions

Aug 2026 · 0 citations · 52 references
Mathematics Computer Science Economics

TL;DR

Findr, short for flexible, interpretable deep regression, is introduced, a semi-structured framework for binary credit risk modelling that decomposes the logit into an interpretable structured component and an orthogonal neural residual.

Abstract

Credit risk models increasingly need to combine predictive accuracy with transparent explanations and auditable fairness constraints. Logistic regression remains attractive because its coefficients are easy to interpret, but it can miss nonlinear structure. Flexible models can improve prediction, but their explanations are often post-hoc and may not describe the decision rule itself. We introduce $\texttt{findr}$, short for flexible, interpretable deep regression, a semi-structured framework for binary credit risk modelling that decomposes the logit into an interpretable structured component and an orthogonal neural residual. The orthogonalisation separates coefficient-based effects from residual nonlinear variation, while an in-processing Wasserstein penalty mitigates group disparities by comparing score distributions during training. The framework also includes diagnostics that measure the structured component's contribution to logit variation, decision agreement, and local directional consistency. We evaluate $\texttt{findr}$ in a simulation study and on eight public credit datasets using score-level accuracy-fairness frontiers. The results show that $\texttt{findr}$ behaves close to logistic regression when the signal is approximately linear, while recovering much of the predictive gain of neural models when nonlinear structure is relevant. The diagnostics identify when coefficient-based explanations remain close to the full fitted model and when residual variation must also be examined. These findings support semi-structured modelling as a practical way to make performance, fairness, and interpretability trade-offs explicit in credit risk decisions.

View source

Similar papers

#small language model Preprint Aug 2026

Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

This work examines whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives and discusses implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.

Sahab Zandi, Noah Kostesku, Christophe Mues et al. · 0 citations
Preprint Aug 2026

Interpretable hybrid credit scoring for thin-file and underbanked populations

We extend a residual-learning hybrid credit scoring framework (logistic regression scorecard plus a gradient-boosting correction on its residuals, decomposed at each prediction into an interpretability ratio $\rho(x)$ that measures the share attributable to the linear branch) along three axes: an East African empirical instantiation on the Zindi Financial Inclusion in Africa data (Kenya, Rwanda, Tanzania, Uganda); a fairness audit at the granularity of the framework's three interpretability regions; and a thin-file segmentation analysis. On the Taiwan Credit Default benchmark retained for continuity, the calibrated hybrid attains AUC $= 0.776$ ($\Delta\mathrm{AUC} = +0.057$ vs.\ standalone logistic regression, $+0.001$ vs.\ standalone XGBoost), reduces Brier Score by 23\%, and concentrates the highest-default-rate borrowers (69.5\%) in the fully interpretable region. On Zindi, the calibrated hybrid attains AUC $= 0.869$ ($\Delta\mathrm{AUC} = +0.015$ vs.\ LR, $p<0.001$; $-0.004$ vs.\ XGBoost), cuts Brier from $0.158$ to $0.085$ (a 46\% reduction), and replicates the regional routing pattern. The fairness audit detects severe routing into the opaque ML-driven region along socioeconomic axes: rural respondents by 18 percentage points relative to urban, primary-or-less-educated by 32 points relative to secondary-and-above, and Ugandan respondents by 22 points relative to Kenyan, while gender shows essentially no routing disparity. The audit pipeline surfaces subgroup-routing violations that aggregate fairness metrics miss, in a form directly usable by African central-bank supervisors of digital credit.

Belise Kanziga, Yaé U. Gaba, Olivier Kanamugire · 0 citations
Open access Jul 2026

An Interpretability Analysis of Credit Default Prediction Using Random Forest with SHAP and LIME

This study explores the use of Explainable Artificial intelligence techniques to improve the interpretability of credit default prediction and highlights the practical value of explainable machine learning in developing more understandable, trustworthy, and accountable credit risk assessment systems for real-world financial decision-making.

Muskan, B. Sidhu · 0 citations
Preprint Jul 2026

Effort-Centric Fairness in Lending Decisions

Algorithmic credit scoring must satisfy fairness and explanation requirements, yet prevailing predictive-parity criteria assess only outcomes at the decision point. They can therefore overlook whether rejected applicants face unequal burdens in reaching future approval, a phenomenon we call masked inequality. We develop an effort-centric framework that measures an applicant's effort as the minimum weighted cost of feasible changes required to cross the approval boundary. The framework distinguishes feature-independent actions from additive structural shifts that propagate through a causal model and defines parity by comparing average minimum effort across protected groups. We derive tractable local expressions for general differentiable classifiers and exact expressions for logistic regression, embed them in an in-processing fairness objective, and bound changes in portfolio credit risk. The same optimisation yields actionable pathways to approval. Using mortgage data with continuous and discrete features, we find that rejected female applicants require greater effort even when standard predictive-parity criteria are satisfied. Feature-independent regularisation reduces the effort gap by more than 50\% with modest predictive changes. Causal regularisation yields reductions above 90\% at the tested positive penalty weights, but with larger predictive and risk-return trade-offs. Expected and unexpected losses remain broadly stable under feature-independent regularisation and increase under causal regularisation; RAROC declines but remains positive. These results show that effort parity complements predictive fairness by revealing and mitigating hidden barriers to future credit access while making the associated operational trade-offs explicit.

Shiqi Fang, Zexun Chen, J. Ansell · 0 citations
Conference Jul 2026

Designing Policy-Compliant Counterfactual Explanations for Fair and Transparent Credit Risk Assessment

The growing adoption of more sophisticated machine learning models in automated decisioning of credit risks has generated very serious issues of explainability, fairness, and consumer trust, especially when loan applications are denied. Alternative methods of explanation that are available like the use of the static reason codes and traditional counterfactual techniques tend to fail to offer realistic, practical and fair advice to the impacted applicants. In this paper, we present a third-generation counterfactual explain model, which combines structural causal modeling, diffusion-based generative learning, fairness-constrained optimization, and policy adaptability control to produce trustworthy and user-friendly credit clarifications. Actionability and real-world consistency are enforced using a structural causal model to separate mutable and immutable attributes and maintain causal relationships between financial variables. A conditional diffusion network is conditioned on approved credit profiles in order to produce several plausible counterfactual representations of applicants. Such candidates are filtered by original credit model to only keep decision-flipping examples to be valid and are optimized over a multi-objective fairness-constrained formulation that balances small feature changes, realism, and diversity, and demographic equity. Additionally, a policy adaptation module, which is based on reinforcement learning, constantly balances the explanation strategy according to the changing lending policies and regulatory issues. The causal diffusion-based framework proposed had greater counterfactual validity, realism, diversity, and fairness as compared to current gradient-based and heuristic approaches on all of the tested credit datasets.

Bhuvaneswari U, S. Muthukrishnan, Pankaj Kumar Baid · 0 citations
Preprint Aug 2026

Estimating the Conditional Forecast-Revision Scale in Sequential Models: Local-Smoothing Limits, Matched Models, and Cost--Accuracy Trade-offs

The \emph{conditional forecast-revision scale} $\It=\{\Var(\E[X_{t+1}\mid\F_t]\mid\F_{t-1})\}^{1/2}$ measures the history-specific size of the forecast update induced by observing $X_t$. Because it is a conditional second moment built from two unknown conditional means, it is not directly observed. We study which estimator of $\It$ should be used under different structural assumptions and computational budgets. The comparison includes a block bootstrap, a conditional-variance model, a fitted state-space model, two $O(1)$ streaming smoothers, and the forget gate of an already-trained recurrent network. An error decomposition separates one-step-prediction error from conditional-second-moment tracking error. We show that externally tuned lag-only smoothers can be inconsistent when $\It$ changes at the sampling scale, although they attain the usual $T^{-2/3}$ mean-squared-error rate ($T^{-1/3}$ for $\It$) under slow variation; a correctly specified state-space estimator escapes this limit by using the current state. In volatility-driven designs, a cheap conditional-variance model is more accurate and over one hundred times cheaper \emph{as a point estimator} than the implemented block bootstrap, whose value lies in the sampling distribution it provides rather than in point tracking. In state-driven designs, only the structurally matched filter recovers the fast variation. Read directly, a trained network's forget gate does not track $\It$ --- though a supervised linear probe on the full gate vector does, so $\It$ is linearly decodable but not available for free. These results yield a practical rule: identify the conditional-second-moment structure, match the estimator to it, and then choose the least costly adequate method.

H. Foo, Y. Chang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.