Credit Default Prediction Using Large Language Models and Machine Learning: An Application to Colombia’s Solidarity Sector
Abstract
Credit default prediction is a standard risk-management task, and large language models (LLMs) have been proposed as prompt-based alternatives, without task-specific parameter updating, for institutions that cannot deploy full machine learning (ML) pipelines. This study evaluates the Informed GPT on Colombian solidarity-sector cooperative lending data, benchmarking gpt-4o-mini against five tuned tree ensemble and gradient boosting classifiers on native imbalanced data (17.3% default rate, 12,861 loans). Six additions relative to the seminal reference are reported: (i) a new empirical domain (Colombian solidarity-sector cooperatives regulated by the SES); (ii) a multi-model benchmark rather than a logistic-regression-only baseline; (iii) a leakage-mitigation prompt design that excludes supervised-analysis-derived hints, causal directions, and target-distribution disclosures; (iv) a calibration analysis using Brier score, log loss, expected calibration error (ECE), reliability diagrams, and calibration slope and intercept; (v) bootstrap 95% confidence intervals, DeLong tests, and McNemar tests for paired significance; and (vi) matched label-budget learning curves for logistic regression, XGBoost, and LightGBM. Tuned ML models attain AUC ≈0.96 (bootstrap CI [0.94,0.98]), while the LLM operates in the AUC 0.67–0.74 range across few-shot sizes N∈{0,10,20,40,80}. Under matched budgets, the LLM outperforms logistic regression at every N but is surpassed by gradient boosting once training samples reach approximately 40–80 observations. LLM probabilities are miscalibrated (ECE 0.11–0.21 vs. ≈0.03 for ML) and over-predict default (mean predicted probability 0.28–0.38 vs. observed base rate 0.17); threshold optimisation and post hoc calibration (Platt scaling, isotonic regression) are required for operational use. The findings qualify earlier claims about LLM auditability and position the approach as an assessment tool for cooperatives with fewer than ≈100 labelled defaults, rather than as a substitute for a well-resourced ML pipeline.