Evidence-Calibrated Financial Language Models for Macro-Policy Stance Classification and Decision Cards
Abstract
Central-bank language often conveys policy direction through qualifying clauses rather than isolated sentiment terms. This study evaluates hawkish, dovish, and neutral stance classification on all 496 FinBen-FOMC excerpts using leakage-controlled five-fold stratified-group cross-validation. Seven systems span a class-prior baseline, an n-gram language model, word-and-character TF-IDF, frozen FinBERT transfer scores, frozen MiniLM sentence embeddings, and two fused models. Four-fold inner cross-validation fits temperature and isotonic calibrators without access to outer-test labels. Evaluation combines macro-F1, Matthews correlation coefficient, expected calibration error, Brier score, confusion analysis, clustered bootstrap intervals, selective risk, and contrastive lexical evidence. TFIDF-LR achieved the highest pooled point estimates for macro-F1 (0.493) and MCC (0.246), followed closely by EvidenceFusion at 0.489 and 0.233. None of the macro-F1 differences between TFIDF-LR and another learned model was significant after Holm correction. Temperature scaling sharply reduced overconfidence in the n-gram and embedding systems; TFIDF+MiniLM attained the lowest temperature-scaled ECE (0.014), while TFIDF-LR achieved a Brier score of 0.587. Isotonic calibration lowered probability loss further for several models but reduced directional recall. Removing TFIDF-LR’s selected evidence terms reduced predicted-class confidence by 0.170 and changed 69.4% of labels, whereas matched random removal changed confidence by −0.019. At a 0.70 acceptance threshold, TFIDF-LR covered 6.45% of records with 21.9% risk, including a high-confidence error driven by tightening vocabulary despite explicit negation. Evidence-calibrated decision cards are therefore most appropriate for conservative analyst triage rather than autonomous policy interpretation.