Sep 2026· Journal of Visualized Experiments· Vol 235· 0 citations
Medicine
TL;DR
A lightweight two-step claim verification framework that decomposes LLM responses into atomic factual claims and independently verifies each extracted claim against a separately generated reference produced through an isolated factual recall prompt, showing consistent performance across the evaluated benchmarks without requiring model retraining.
Abstract
Large Language Models (LLMs) exhibit a critical tendency to generate factually incorrect yet linguistically fluent outputs - a phenomenon termed hallucination - which poses serious risks in precision-critical applications. Existing mitigation strategies, including retrieval-augmented generation and self-consistency sampling, either introduce substantial inference latency or depend on external knowledge infrastructure, limiting their applicability in real-time deployments. This paper proposes a lightweight two-step claim verification framework that decomposes LLM responses into atomic factual claims and independently verifies each extracted claim against a separately generated reference produced through an isolated factual recall prompt. Although the generator and verifier share the same underlying language model, separating response generation from factual recall reduces direct response conditioning and mitigates confirmation bias during verification, using Natural Language Inference, and applies an adaptive statistical threshold - defined as τ = µ + kσ over the NLI confidence score distribution - to selectively correct only contradicted claims. Unlike prior NLI-based methods that rely on fixed decision boundaries, the proposed framework dynamically adapts its verification threshold to the confidence distribution of each response, showing consistent performance across the evaluated benchmarks without requiring model retraining. Evaluated on TruthfulQA and FEVER, the framework reduces the hallucination rate from 28% to 9% on TruthfulQA - a 67.9% relative reduction - while incurring only 160 ms of additional latency over the baseline LLM and outperforming SelfCheckGPT and FActScore in hallucination detection accuracy. These results indicate that the framework can provide a favorable balance between factual reliability and response latency on the evaluated benchmarks, while further validation across domains and deployment settings is needed.
A self-reflective framework in which an LLM generates an answer, identifies claims that may be uncertain, performs an internal verification stage, and revises the response before delivery is proposed.
Priti Sharma, Sachin Sharma· Iconic research and engineer...· 0 citations
Large language models (LLMs) can generate fluent and confident responses that are factually incorrect, unsupported by evidence, or inconsistent with the source material. These hallucinations reduce the reliability of LLM-based question answering, summariza tion, dialogue, and information retrieval systems, especially w...
CAL-RAG (Context-Aware Low-Rank Calibration for RAG), a parameter-efficient fine-tuning and decoding calibration framework designed to enforce strict contextual faithfulness without compromising generative fluency, is proposed.
Sophia N. Tawar, Liam K. Peing, Amani Bellow· International Journal of App...· 0 citations
Large Language Models (LLMs) can generate fluent and convincing responses, but fluency does not guarantee
factual correctness. Hallucination occurs when a model produces information that is false, unsupported, or inconsistent
with available evidence. This paper reviews why hallucinations arise andexamine Retrieval-Augm...
Shyalaja L. N., Shantinath Patil, P. R. et al.· International Journal for Re...· 0 citations
Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and temperature-scaled calibration for response-level ha...
Large language models (LLMs) have come into widespread use in recent years across domains ranging from education and journalism to academic research and everyday information seeking. Their ability to produce fluent, coherent-sounding answers in natural language creates the impression that these answers are accurate and...
Harun Nabiyev· Aposta: Revista de Ciencias...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.