This work introduces TempFinRAG, a point-in-time evaluation protocol built from public filings and XBRL facts, and introduces TempFinQA, a point-in-time evaluation protocol built from public filings and XBRL facts, and evaluates the framework on complementary evidence-grounded, numerical, conversational, and multi-table benchmarks.
Abstract
Financial question answering is often treated as document question answering, although financial evidence is both multimodal and time-dependent. Semantically equivalent facts expressed in narrative text, tables, page images, or Extensible Business Reporting Language (XBRL) should support consistent answers, whereas a disclosure may support a query only after becoming public. We formalise this combination as crossmodal evidence symmetry under a causal temporal boundary and introduce TempFinRAG, a multimodal temporal retrieval-augmented generation (RAG) framework for point-in-time financial question answering. Given a question, company, and as-of date, the framework enforces the information boundary defined by U.S. Securities and Exchange Commission (SEC) filing availability; aligns page layout, text, table structure, and XBRL facts; retrieves time-valid evidence; executes auditable financial calculations; and generates a cited answer. A verifier checks temporal validity, claim support, numerical consistency, and the need to abstain. We further introduce TempFinQA, a point-in-time evaluation protocol built from public filings and XBRL facts, and evaluate the framework on complementary evidence-grounded, numerical, conversational, and multi-table benchmarks. On TempFinQA, TempFinRAG improves answer accuracy from 66.7% to 78.9% over hybrid RAG while reducing temporal evidence leakage from 10.8% to 1.7% and hallucination from 17.3% to 7.9%. Reliable financial question answering therefore requires consistent treatment across evidence representations and deliberately asymmetric access across time.
TimelyRAG is proposed, a retriever-agnostic framework that incorporates temporal distance into ranking to align queries with version-appropriate documents, and TimelyQABench is introduced, the first benchmark for regulation-heavy domains with overlapping-evolving challenges.
Youngeun Nam, Joeun Kim, Hwanjun Song et al.· 1 citation
This work studies whether a carefully domain-adapted retrieval-augmented generation pipeline closes the gap between compact and compact model quality in financial institutions under dense, frequently amended rulebooks.
Tobias Deußer, Abhishek Pillai, A. Bariviera et al.· 0 citations
Results align with a diagnostic perspective on chunking: using evidence at a task-appropriate level of granularity can improve grounding, auditability, and answer quality, but the observed patterns should be interpreted within the HotpotQA distractor setting, fixed generator, and tested context budgets.
: Event knowledge graphs support event-centric question answering by linking events to temporal, location, participant, and source-record information, but semantic relevance alone does not guarantee that a selected event satisfies every represented condition. We propose EviGraphRAG, a ranker-agnostic reliability layer...
Yu-Teng Sun, Yang Su, Xu-An Wang· Computers, Materials & C...· 0 citations
By evaluating paper retrieval, evidence grounding, and answer accuracy separately, LitTraceQA provides a testbed for scientific QA systems that produce verifiable answers rather than unsupported summaries.
FinRAG-QA is introduced, a novel benchmark dataset for financial question answering, which comprises 999 practitioner-curated questions on 10 standardised indicators, grounded in 209 annual and Pillar 3 reports from 24 major European and U.S. banks spanning 2019-2023.
Arianna Miola, Bruno Spaccavento, Lorenzo Silotto et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.