Skip to content

FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

Sep 2026 · 0 citations · 40 references
Computer Science

TL;DR

FinFIRST is the first financial benchmark to jointly evaluate answers and supporting evidence through atomic rubrics, retaining final-answer correctness as the primary objective while making the supporting research process measurable, verifiable, and diagnosable.

Abstract

Financial search is a highly demanding task for LLM agents, requiring not only a correct final answer but also temporally valid information retrieval, authoritative source selection, entity and period alignment, unit and definition consistency, and verifiable evidence for all conclusions. Existing benchmarks predominantly evaluate only the final answer, making it difficult to localize errors or assess whether an answer is well-founded. To address this gap, we introduce FinFIRST (Financial Information Retrieval, Sourcing and Traceability), the first financial benchmark to jointly evaluate answers and supporting evidence through atomic rubrics. FinFIRST comprises 123 expert-authored tasks spanning a graduated difficulty spectrum, constructed from aggregate patterns of real-world financial scenarios through an 18-field taxonomy, a six-axis coverage blueprint, a registry of 138 financial sources, contributions from over 50 finance experts, and a six-stage quality-control pipeline. Each task is accompanied by an evidence-grounded reference package decomposed into atomic criteria across three dimensions: raw-information acquisition, source verification, and computation and answer formation. We evaluate 15 model configurations under a unified tool setting. Claude-Opus-5 achieves the highest atomic score of 87.59%, while GPT-5.6-Sol attains the highest strict pass rate of 71.54%. Computation and answer formation consistently lag behind raw-information acquisition across systems. FinFIRST retains final-answer correctness as the primary objective while making the supporting research process measurable, verifiable, and diagnosable.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase orders, approvals, allocations, payments, l...

Prof. S. B. Ghawate · 1 citation
Preprint Aug 2026

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

CorporateBench (CB), a human-validated multi-task Q&A benchmark whose scale approaches the conditions LLMs encounter in corporate communication networks, is presented, with evaluation corpora surpassing 230,000 documents.

Sil Hamilton, Albert Yu Sun, Óscar Romero et al. · 0 citations
#natural language process... Preprint Sep 2026

TRACE: Target-Aware Retrieval, Attributed Evidence, and Contract-Constrained Extraction for LitTraceQA

Finding a relevant paper is not the same as producing a verifiable answer from it. LitTraceQA requires canonical paper identifiers, exact evidence at the page or object level, and typed answers that match the evaluator. We call the separation between source access and scorer-visible correctness the grounding contract g...

Sachin Gupta, Divya Godara · 0 citations
#artificial intelligence Preprint Sep 2026

DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents

Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical computation, and constructing such benchmarks manually is costly. We present DI-Bench, a pipeline for generating realistic benchmarks for data int...

Jiang-Yun Zhang, K. Surrao, Torpong Nitayanont et al. · 0 citations
Preprint Sep 2026

KnowFeat: Knowledge-Guided Feature Engineering via LLM Agents

Automated feature engineering with large language models (LLMs) can produce semantically meaningful features for tabular data, yet existing methods lack structured domain knowledge, rigorous verification, and explainable provenance. We propose KnowFeat, a knowledge-guided feature engineering framework that organizes do...

Chengsong You, Wangyue Li, Wei-Qiao Que et al. · 0 citations
#artificial intelligence Preprint Sep 2026

BudgetVerify: Budget-Tiered Verification for Financial QA

Financial question answering often requires precise numerical extraction, unit handling, and arithmetic over tables and text, but applying expensive verification uniformly wastes test-time compute. We propose BudgetVerify, a budget-tiered generator-verifier framework that routes each generated answer to one of three ve...

Janet Jenq, Hong-Da Shen · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.