Skip to content
Review Open access

Ensuring Data Integrity in Official Financial Statistics: A Review of Hybrid AI and XAI Methods in the Context of Market Efficiency and Value Investing

Jul 2026 · Journal of Official Statistics · 0 citations · 29 references

Abstract

The integrity and accuracy of financial data are prerequisites for market efficiency; however, data anomalies and quality issues severely compromise their “fitness for use” in sophisticated decision-making processes, such as value investing strategies. This article reviews the application of advanced artificial intelligence (AI) methods to enhance quality assurance, anomaly detection, and imputation within high-dimensional financial data streams. The paper critically evaluates both statistical-machine learning hybrids (e.g., ARIMA-LSTM) and deep learning combinations (e.g., autoencoder-based GANs), alongside Explainable Artificial Intelligence (XAI) techniques, assessing their utility against the strict auditability requirements of public trust institutions. The synthesized literature suggests that hybrid frameworks can potentially outperform monolithic approaches in detecting nonlinear manipulations and creating “high-fidelity” datasets. Furthermore, the study addresses the “black box” opacity challenge—a major barrier for regulatory and statistical agencies—discussing how methods like SHAP and LIME support, rather than independently ensure, the necessary interpretability of algorithmic decisions. Conclusions indicate that the synergy between the predictive power of advanced AI models and the transparency supported by XAI is a highly valuable component for modern market supervision, enabling effective data validation while supporting institutional accountability.

Read PDF

Similar papers

Open access Jul 2026

From Black Box to Boardroom: The Significance of Explainable AI (XAI) in Reducing Algorithmic Risk and Rebuilding Confidence in Digital Payment Systems

The wide implementation of advanced Machine Learning (ML) models in digital payment systems, especially for fraud detection and credit risk assessment, has substantially improved operational efficiency and transaction security. The inherent opacity, often referred to as the black box character, of these high-performing algorithms poses considerable and mounting issues related to algorithmic fairness, stakeholder trust, and compliance with regulations. This article analyzes the growing strategic significance of Explainable Artificial Intelligence (XAI) as an important governance tool for mitigating algorithmic risk in financial services. The paper exposes how XAI, informed by Agency Theory and Institutional Theory, is not just a technical requirement but an essential institutional mechanism for ensuring regulatory accountability within frameworks like the EU AI Act, restoring public trust and identifying and alleviating systemic algorithmic bias in credit scoring and fraud risk assessment. A conceptual framework is introduced and it illustrates how XAI; using post-hoc interpretation methods such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations)- bridges the knowledge disparity between intricate AI models and various human stakeholders, including customers, fraud analysts, and regulators. This transformation shifts AI from a hypothetical institutional liability to a responsible, auditable, and governable asset within the digital payment ecosystem. The report concluded by describing key areas for forthcoming empirical research on the organizational problems associated with XAI implementation across various regulatory jurisdictions

Temitope Onibaniyi, Umar Lawal · 0 citations
Review Jul 2026

A Review of Machine Learning Applications for Credit Default Risk Prediction and Early Warning Systems

The paper shows that ensemble learning models have superior predictive power and argues that behavioral data complement traditional datasets for underbanked populations, such as "credit invisibles," and makes a strong case for XAI being essential for model transparency, combating bias, and meeting regulatory requirements.

Bojun Chen · 0 citations
#explainable ai Review Open access Aug 2026

Data analytics and internal controls in U.S. financial systems: A review of fraud detection and financial reporting integrity

The study concludes that strengthening fraud detection and financial reporting integrity requires integrating analytics and internal controls within a unified governance framework supported by continuous monitoring, institutional accountability, and transparent oversight mechanisms.

Francesca Nyarkoa Kobla, Jessica Fosua Agyei · 0 citations
Review Open access Aug 2026

Machine Learning for Individual Credit Risk Assessment: A Systematic Literature Review of State-of-the-Art Methods, Challenges and Perspectives

Credit risk assessment forms a cornerstone of banking risk management and the stability of the wider financial system. Over the past decade, the rapid development of machine learning (ML) techniques has substantially enhanced traditional credit risk assessment methodologies. ML has now emerged as a core technological pillar for the banking sector, strengthening risk identification capabilities, optimising credit decision-making, and advancing financial inclusion. Conventional credit scoring models, dominated by logistic regression (LR) and scorecard approaches, offer inherent strengths in interpretability and regulatory compliance. However, constrained by their linear assumptions, these methods struggle to capture complex non-linear relationships within credit data and deliver insufficient predictive accuracy for the “credit-invisible” population lacking formal credit histories. This paper presents a systematic literature review (SLR) of ML applications in credit risk assessment (CRA), covering publications from January 2016 to May 2026. A total of 894 papers were retrieved from five digital libraries, and following a rigorous multi-stage screening process, 129 studies were selected for final inclusion. Our analysis reveals that tree-based ensemble models and deep learning (DL) architectures predominate in contemporary research in this field. Meanwhile, post hoc explanation methods and machine learning operations (MLOps) are gaining significant traction as solutions to address fairness, transparency, and system maintenance challenges in real-world production environments. We synthesise prevailing methodologies into a unified end-to-end credit risk modelling framework spanning data preprocessing, feature engineering, model training, evaluation, and operational deployment. Through a critical assessment of the advantages, limitations, and inherent trade-offs of existing approaches, this SLR not only identifies current research gaps and future directions for the academic community, but also provides practical guidance for the banking sector to build compliant, fair, and efficient intelligent risk assessment systems.

Bolun Zhang, Jun Luo, Ruobing Wu et al. · 0 citations
Preprint Jul 2026

Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks

Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and under-utilized textual data in financial reports. Existing methods often rely on random data splits, leading to overoptimistic performance estimates that do not reflect real-world generalization to new companies or future periods. To address this recurring problem with the state of the art, we propose a robust FSFD framework leveraging Large Language Models (LLMs) to integrate both structured financial data and unstructured textual information from financial reports. We provide a more realistic evaluation through a novel and challenging benchmark task called Company-Isolated FSFD (CI-FSFD). We construct and make publicly available a comprehensive U.S. company dataset combining financial statements, summarized MD&A text, and fraud labels. Our approach achieves the best performance on the challenging CI-FSFD task, demonstrating the critical value of textual data and robust evaluation for reliable financial fraud detection.

Guy Stephane Waffo Dzuyo, Gaël Guibon, Christophe Cerisara et al. · 0 citations
Review Open access Jul 2026

The Role of Statistical Concepts in the Development of Artificial Intelligence for Auditing: A Systematic Literature Review

Artificial Intelligence (AI) has increasingly been adopted in auditing, however the role of statistical methods in improving AI performance and reliability remains fragmented. This study aims to systematically identify, review, and synthesize the role of statistical concepts in AI development for auditing, including their implementation, benefits, challenges, and future directions. A Systematic Literature Review (SLR) following the PRISMA framework was conducted on 149 studies published between 2017 and 2026 from Google Scholar, Scopus, and SciSpace. The findings reveal a significant increase in AI auditing research since 2023. Regression analysis, hypothesis testing, and Bayesian inference are the most frequently applied statistical methods, supporting fraud detection, risk assessment, audit sampling, anomaly detection, and model validation. Integrating statistical methods with AI improves prediction accuracy, interpretability, transparency, and audit quality. However, challenges remain regarding data quality, model interpretability, auditor competency, and AI governance. Future research should prioritize hybrid AI-statistical models, Explainable AI, and adaptive Bayesian approaches to enhance trustworthy data-driven auditing

Nibrisatul Hana, Shofyan Hadi, Lintang Budiarti et al. · 0 citations