Jul 2026· Annual International Computer Software and Applications Conference· pp. 1611-1620· 0 citations· 35 references
Computer Science
Abstract
Financial regulatory documents are characterized by their fine-grained complexity and pronounced heterogeneity, featuring specialized domain-specific content and diverse structural formats that vary across different regulatory frameworks and jurisdictions. These characteristics challenge modern Question-Answering (QA) systems, which often suffer from limited domain adaptation, poor interpretability, and hallucinatory problems. This work was conducted within a banking and software company, where such challenges directly impact regulatory compliance efforts. Our goal is to introduce RegulQA, a hybrid QA system that can extract accurate, logical, and comprehensible answers from unstructured regulatory documents. RegulQA integrates knowledge graph reasoning, semantic search, and retrieval-augmented generation using large language models. Experimental evaluation shows that RegulQA improves QA performance and significantly reduces hallucination rates. The baseline model employing only the LLM demonstrated a hallucination rate of 24%, whereas the proposed approach, combining knowledge graph reasoning and semantic retrieval with LLM reasoning, effectively reduced the hallucination rate to approximately the half. This integrated approach also yielded the best balanced overall scores across key qualitative attributes, including coverage, non-redundancy, readability, and response quality.
This work introduces RegNLI, a novel framework that formulates misbranding detection as a inference task between product claims and regulatory provisions, and builds a foundation for compliance-aware NLP systems and opens new directions for integrating formal reasoning with neural architectures in regulatory domains.
Findings identify evidence localization and fine-grained type discrimination as distinct challenges and show that compact supervised encoders are strong baselines for this task.
Aman Kumar, Lasitha Vidyaratne, Dipanjan Ghosh et al.· arXiv.org· 0 citations
Analyzing financial documents such as 10-K filings, tabular disclosures, and macroeconomic reports demands expert reasoning and extensive time. However, existing Retrieval-Augmented Generation systems often struggle to process hybrid text-table structures or the massive scale of financial documents. To address these ch...
Debate-on-Graph (DoG) is proposed, a new framework that enables LLMs and UKGs to collaborate adaptively for reliable reasoning and introduces a Multi-Agent Debate mechanism, which yields reliable answers through adaptive adversarial debates, aiming to fully exploit the knowledge in UKGs while preserving the reliability...
Financial text, textbooks, and question-answer pairs are abundant, but only a small fraction is directly usable for reasoning-focused post-training. Existing QA pairs often lack explicit reasoning, sufficient context, or reliably verifiable answers, while textbooks must first be transformed into synthetic training exam...
Zhirayr Hayrapetyan, Andrei Kalmykov, Denis V. Kokosinskii et al.· 0 citations
This work proposes a comprehensive pipeline for improving financial QA systems through high-quality synthetic data generation and fine-tuning of smaller language models (SLMs) using Quantized Low-Rank Adaptation (QLoRA).
Lokendra Birla, Milind Savagaonkar, Visnu Srinivasan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.