Skip to content

Measurement Risk in LLM-Based Financial NLP: Rubric and Metric Sensitivity on JF-ICR

Sep 2026 · IEEE Conference on Computational Intelligence for Financial Engineering & Economics · pp. 230-235 · 0 citations · 18 references

Abstract

Large language models are increasingly used to read earnings calls, investor-relations Q&A, guidance, and disclosure language. In this setting, supervised financial NLP benchmarks can become evidence for vendor selection, deployment approval, and model-risk records. Gold labels, however, do not make a benchmark score a neutral fact. We study measurement risk in supervised financial NLP: the risk that the evaluation instrument, rather than the model alone, determines the reported result. The case study is Japanese Financial Implicit-Commitment Recognition (JF-ICR), a five-class ordinal financial Q&A task. We audit a pinned 253-item test split across four commercial LLM classifiers, five rubrics, three temperatures, five ordinal metrics, and three rank aggregators. Three findings follow. First, rubric wording materially changes item-level labels: literal versus pragmatic rubrics agree on 70.0–83.4% of labels, with movement concentrated near the +1/0 boundary. This pattern is consistent with a pragmatic-boundary interpretation, but not a validated linguistic-causality claim because the rubric variants confound semantics, examples, and verbosity. Second, some metrics are weakly identified under this class distribution. Within-one accuracy is too easy because near misses receive credit and the majority class dominates; worst-class accuracy is too noisy because the rarest class has only two examples. Exact accuracy, macro-F1, and weighted kappa are more interpretable for primary ranking under our operational rule. Third, ranking claims become more defensible only after this metric-identifiability audit: Bradley–Terry, Borda, and Ranked Pairs agree on the identifiable metric subset, while the full five-metric sweep creates a close-pair disagreement. The contribution is a reporting discipline for financial NLP benchmarks: before scores guide model selection or deployment, the rubric, metric, aggregation policy, and benchmark provenance should be audited.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.