Skip to content
Open access

An Evidence-Driven Hybrid Architecture for Financial Statement Fraud Detection With Long-Context Transformers

2026 · IEEE Access · Vol 14, pp. 115211-115230 · 0 citations · 23 references
Computer Science

TL;DR

In this benchmark, hierarchical encoders offer the strongest single-model accuracy–efficiency trade-off, outperforming the sparse-attention baseline at lower memory and latency, while head+tail BERT remains a competitive low-cost anchor.

Abstract

Detecting financial statement fraud from narrative disclosures is difficult because evidence is sparse and distributed across long, structured public-company annual reports filed on Form 10-K. We present a time-aware benchmark focused on Management’s Discussion and Analysis (MD&A) for financial statement fraud detection (FSFD) and evaluate representative long-document architectures under time-aware forward-chaining splits. We compare hierarchical encoders (hierarchical attention transformer and hierarchical document transformer), a finance-domain long encoder (LongFinBERT), a sparse-attention model (Longformer/LED), and a short-context head+tail BERT baseline. Performance is assessed using the area under the receiver operating characteristic curve (AUROC), average precision, audit-budget ranking quality based on cumulative gain, and deployability metrics such as peak GPU memory and per-filing latency. We further propose an evidence-driven Hybrid pipeline (Selector → Reranker → Calibrated Stacker) that selects paragraph evidence with high recall, applies cross-evidence reranking, and fuses calibrated scores with a monotone stacker. On the held-out 2014–2019 horizon, the Hybrid achieves AUROC = 0.871 and average precision = 0.392 and improves audit-budget ranking over the strongest single model. In our benchmark, hierarchical encoders offer the strongest single-model accuracy–efficiency trade-off, outperforming the sparse-attention baseline at lower memory and latency, while head+tail BERT remains a competitive low-cost anchor. Year- and issuer-stratified block bootstrap analysis indicates that gains in average precision and cumulative-gain ranking quality are robust under our time-aware evaluation protocol.

Read PDF

Similar papers

Open access Jul 2026

Next-Generation Payment Fraud Intelligence with Large Language Models and Event-Driven Architectures

Payment fraud detection at scale must reconcile three pressures: the streaming nature of transaction events, the cost-sensitivity of false alerts, and the practical need for models that can be trained and refreshed on commodity hardware. Recent work has explored advanced text representation models and event-driven arch...

Bingjie Zi · 0 citations
#artificial intelligence Preprint Sep 2026

Graph-Transformer Fraud Detection with Self-Supervised Pretraining and Conformal Risk Control

Financial fraud in corporate transaction networks has grown more coordinated and harder to detect with rule-based engines and with classical learning models that treat each transaction in isolation. This paper presents GTFD, a graph-transformer fraud detector that fuses structural and temporal evidence from a corporati...

Sergei, Komarov · 0 citations
#graph neural networks Open access Sep 2026

Design of a financial fraud detection model optimized by multi-task learning and graph neural networks

Contemporary financial regulation and risk identification are increasingly challenged by the escalating intricacy of inter-firm relational architectures, the diversification of financial conduct, and the multidimensionality of data sources. Conventional fraud detection methodologies, predominantly grounded in single-ta...

Ding-Mou Huang, Lian Hu, Muhammad Asif · 0 citations
Open access Aug 2026

Deep Learning Framework for Financial Fraud Detection: Systematic Feature Engineering and Comparative Evaluation of Neural Architectures

Three deep tabular models, namely, an advanced multilayer perceptron (AdvancedMLP), an attention‐based residual network (AttentionFraudNet), and an advanced residual network (AdvancedResNet), are compared against three traditional machine learning baselines, including Random Forest, Gradient Boosting, and Logistic Regr...

Vahid Azarvand, Parvin Azhdari, A. Beitollahi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.