Measurement Risk in LLM-Based Financial NLP: Rubric and Metric Sensitivity on JF-ICR
Abstract
Large language models are increasingly used to read earnings calls, investor-relations Q&A, guidance, and disclosure language. In this setting, supervised financial NLP benchmarks can become evidence for vendor selection, deployment approval, and model-risk records. Gold labels, however, do not make a benchmark score a neutral fact. We study measurement risk in supervised financial NLP: the risk that the evaluation instrument, rather than the model alone, determines the reported result. The case study is Japanese Financial Implicit-Commitment Recognition (JF-ICR), a five-class ordinal financial Q&A task. We audit a pinned 253-item test split across four commercial LLM classifiers, five rubrics, three temperatures, five ordinal metrics, and three rank aggregators. Three findings follow. First, rubric wording materially changes item-level labels: literal versus pragmatic rubrics agree on 70.0–83.4% of labels, with movement concentrated near the +1/0 boundary. This pattern is consistent with a pragmatic-boundary interpretation, but not a validated linguistic-causality claim because the rubric variants confound semantics, examples, and verbosity. Second, some metrics are weakly identified under this class distribution. Within-one accuracy is too easy because near misses receive credit and the majority class dominates; worst-class accuracy is too noisy because the rarest class has only two examples. Exact accuracy, macro-F1, and weighted kappa are more interpretable for primary ranking under our operational rule. Third, ranking claims become more defensible only after this metric-identifiability audit: Bradley–Terry, Borda, and Ranked Pairs agree on the identifiable metric subset, while the full five-metric sweep creates a close-pair disagreement. The contribution is a reporting discipline for financial NLP benchmarks: before scores guide model selection or deployment, the rubric, metric, aggregation policy, and benchmark provenance should be audited.