Skip to content

Adaptive NLI-Driven Claim Verification with Statistical Decision Modeling for Low-Latency Hallucination Reduction in Large Language Models.

Sep 2026 · Journal of Visualized Experiments · Vol 235 · 0 citations
Medicine

TL;DR

A lightweight two-step claim verification framework that decomposes LLM responses into atomic factual claims and independently verifies each extracted claim against a separately generated reference produced through an isolated factual recall prompt, showing consistent performance across the evaluated benchmarks without requiring model retraining.

Abstract

Large Language Models (LLMs) exhibit a critical tendency to generate factually incorrect yet linguistically fluent outputs - a phenomenon termed hallucination - which poses serious risks in precision-critical applications. Existing mitigation strategies, including retrieval-augmented generation and self-consistency sampling, either introduce substantial inference latency or depend on external knowledge infrastructure, limiting their applicability in real-time deployments. This paper proposes a lightweight two-step claim verification framework that decomposes LLM responses into atomic factual claims and independently verifies each extracted claim against a separately generated reference produced through an isolated factual recall prompt. Although the generator and verifier share the same underlying language model, separating response generation from factual recall reduces direct response conditioning and mitigates confirmation bias during verification, using Natural Language Inference, and applies an adaptive statistical threshold - defined as τ = µ + kσ over the NLI confidence score distribution - to selectively correct only contradicted claims. Unlike prior NLI-based methods that rely on fixed decision boundaries, the proposed framework dynamically adapts its verification threshold to the confidence distribution of each response, showing consistent performance across the evaluated benchmarks without requiring model retraining. Evaluated on TruthfulQA and FEVER, the framework reduces the hallucination rate from 28% to 9% on TruthfulQA - a 67.9% relative reduction - while incurring only 160 ms of additional latency over the baseline LLM and outperforming SelfCheckGPT and FActScore in hallucination detection accuracy. These results indicate that the framework can provide a favorable balance between factual reliability and response latency on the evaluated benchmarks, while further validation across domains and deployment settings is needed.

View source

Similar papers

Conference Open access Sep 2026

Mitigating Hallucinations in Natural Language Generation through Prompt Engineering: A Mechanism- Oriented Narrative Review

Large language models (LLMs) can generate fluent and confident responses that are factually incorrect, unsupported by evidence, or inconsistent with the source material. These hallucinations reduce the reliability of LLM-based question answering, summariza tion, dialogue, and information retrieval systems, especially w...

Jia-Nian Lin · 0 citations
Open access Aug 2026

Parameter-Efficient Contextual Calibration for Hallucination Mitigation in Domain-Specific Large Language Model Retrieval-Augmented Generation

CAL-RAG (Context-Aware Low-Rank Calibration for RAG), a parameter-efficient fine-tuning and decoding calibration framework designed to enforce strict contextual faithfulness without compromising generative fluency, is proposed.

Sophia N. Tawar, Liam K. Peing, Amani Bellow · 0 citations
Review Open access Aug 2026

Large Language Models Hallucinate and How Retrieval- Augmented Generation Mitigates It

Large Language Models (LLMs) can generate fluent and convincing responses, but fluency does not guarantee factual correctness. Hallucination occurs when a model produces information that is false, unsupported, or inconsistent with available evidence. This paper reviews why hallucinations arise andexamine Retrieval-Augm...

Shyalaja L. N., Shantinath Patil, P. R. et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Domain-Specific Hallucination Detection in Large Language Models

Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and temperature-scaled calibration for response-level ha...

Varun Teja Chundru, Debasmita Biswas · 0 citations
Review Open access Sep 2026

Information retrieval reliability in large language models: a study of source verification

Large language models (LLMs) have come into widespread use in recent years across domains ranging from education and journalism to academic research and everyday information seeking. Their ability to produce fluent, coherent-sounding answers in natural language creates the impression that these answers are accurate and...

Harun Nabiyev · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.