In this benchmark, hierarchical encoders offer the strongest single-model accuracy–efficiency trade-off, outperforming the sparse-attention baseline at lower memory and latency, while head+tail BERT remains a competitive low-cost anchor.
Abstract
Detecting financial statement fraud from narrative disclosures is difficult because evidence is sparse and distributed across long, structured public-company annual reports filed on Form 10-K. We present a time-aware benchmark focused on Management’s Discussion and Analysis (MD&A) for financial statement fraud detection (FSFD) and evaluate representative long-document architectures under time-aware forward-chaining splits. We compare hierarchical encoders (hierarchical attention transformer and hierarchical document transformer), a finance-domain long encoder (LongFinBERT), a sparse-attention model (Longformer/LED), and a short-context head+tail BERT baseline. Performance is assessed using the area under the receiver operating characteristic curve (AUROC), average precision, audit-budget ranking quality based on cumulative gain, and deployability metrics such as peak GPU memory and per-filing latency. We further propose an evidence-driven Hybrid pipeline (Selector → Reranker → Calibrated Stacker) that selects paragraph evidence with high recall, applies cross-evidence reranking, and fuses calibrated scores with a monotone stacker. On the held-out 2014–2019 horizon, the Hybrid achieves AUROC = 0.871 and average precision = 0.392 and improves audit-budget ranking over the strongest single model. In our benchmark, hierarchical encoders offer the strongest single-model accuracy–efficiency trade-off, outperforming the sparse-attention baseline at lower memory and latency, while head+tail BERT remains a competitive low-cost anchor. Year- and issuer-stratified block bootstrap analysis indicates that gains in average precision and cumulative-gain ranking quality are robust under our time-aware evaluation protocol.
Payment fraud detection at scale must reconcile three pressures: the streaming nature of transaction events, the cost-sensitivity of false alerts, and the practical need for models that can be trained and refreshed on commodity hardware. Recent work has explored advanced text representation models and event-driven arch...
A multi-view ensemble ML framework that combines Extreme Gradient Boosting for known fraud patterns, Isolation Forest for label-free anomaly detection, and Graph Sample and Aggregate for relational patterns associated with transaction activities is proposed.
Financial fraud in corporate transaction networks has grown more coordinated and harder to detect with rule-based engines and with classical learning models that treat each transaction in isolation. This paper presents GTFD, a graph-transformer fraud detector that fuses structural and temporal evidence from a corporati...
Contemporary financial regulation and risk identification are increasingly challenged by the escalating intricacy of inter-firm relational architectures, the diversification of financial conduct, and the multidimensionality of data sources. Conventional fraud detection methodologies, predominantly grounded in single-ta...
Ding-Mou Huang, Lian Hu, Muhammad Asif· PeerJ Computer Science· 0 citations
Three deep tabular models, namely, an advanced multilayer perceptron (AdvancedMLP), an attention‐based residual network (AttentionFraudNet), and an advanced residual network (AdvancedResNet), are compared against three traditional machine learning baselines, including Random Forest, Gradient Boosting, and Logistic Regr...
Vahid Azarvand, Parvin Azhdari, A. Beitollahi· Engineering Reports· 0 citations