Aug 2026· International Journal of Applied Engineering and Intelligent Computing· Vol 1, pp. 14-16· 0 citations
TL;DR
CAL-RAG (Context-Aware Low-Rank Calibration for RAG), a parameter-efficient fine-tuning and decoding calibration framework designed to enforce strict contextual faithfulness without compromising generative fluency, is proposed.
Abstract
Retrieval-Augmented Generation (RAG) has become the gold standard paradigm for deploying Large Language Models (LLMs) in knowledge-intensive and high-stakes domains such as biomedical inquiry, financial compliance, and legal reasoning. Despite providing external grounding documents, LLMs continue to exhibit insidious factual hallucinations—either by fabricating plausible-sounding unsupported assertions or by ignoring conflicting retrieved evidence in favor of memorized parametric training biases. Existing mitigation approaches, such as full-parameter fine-tuning or iterative self-reflection prompting, incur prohibitive computational costs and excessive inference latency. In this paper, we propose CAL-RAG (Context-Aware Low-Rank Calibration for RAG), a parameter-efficient fine-tuning and decoding calibration framework designed to enforce strict contextual faithfulness without compromising generative fluency. CAL-RAG introduces a dual-channel Low-Rank Adaptation (LoRA) mechanism: a Context-Grounded Adapter that measures token-level semantic consistency against retrieved evidence chunks, and an Entropy-Gated Decoding Controller that dynamically modulates vocabulary probability distributions during autoregressive generation based on cross-attention dispersion. We conduct extensive empirical evaluations across three challenging domain benchmarks: BioASQ (biomedical), FinQA (financial reasoning), and LegalBench (legal clause interpretation), utilizing open-source LLM backbones (Llama-3-8B, Mistral-7B-Instruct, and Gemma-7B). CAL-RAG reduces factual hallucination rates by 43.7% relative to standard RAG baselines while improving Faithfulness Score from 0.642 to 0.891 and answer accuracy by +11.8% F1. Remarkably, CAL-RAG adds only 0.4% trainable parameters and introduces less than 6.5 ms token latency overhead, making it highly suitable for enterprise production deployment.
Large Language Models (LLMs) can generate fluent and convincing responses, but fluency does not guarantee
factual correctness. Hallucination occurs when a model produces information that is false, unsupported, or inconsistent
with available evidence. This paper reviews why hallucinations arise andexamine Retrieval-Augm...
Shyalaja L. N., Shantinath Patil, P. R. et al.· International Journal for Re...· 0 citations
A lightweight two-step claim verification framework that decomposes LLM responses into atomic factual claims and independently verifies each extracted claim against a separately generated reference produced through an isolated factual recall prompt, showing consistent performance across the evaluated benchmarks without...
Subasish Mohapatra, Biswajeet Dash, Subhadarshini Mohanty et al.· Journal of Visualized Experi...· 0 citations
Retrieval-Augmented Generation (RAG) has been cited as a potential technique for reducing hallucinations in Large Language Models (LLM). However, applying RAG in the Chinese legal landscape remains a challenge due to the large semantic gap between informal user queries and professional legal terminology, and the risk o...
Haoran Li· Mathematical Modeling and Al...· 0 citations
Retrieval-Augmented Generation (RAG) mitigates knowledge obsolescence and factual hallucination in large language models by introducing external context. However, when retrieved knowledge conflicts with the model's internal parametric knowledge, the model may either blindly follow misleading context or incorrectly rely...
Zheng-Chen Huang, Yun-Dong Sun, Min-Rui Song et al.· 0 citations
A self-reflective framework in which an LLM generates an answer, identifies claims that may be uncertain, performs an internal verification stage, and revises the response before delivery is proposed.
Priti Sharma, Sachin Sharma· Iconic research and engineer...· 0 citations
Multimodal large language models (MLLMs) have made strong progress on visual question answering and image captioning, yet they still produce fluent claims about objects, attributes, or relations that are not grounded in the image. Many remedies either modify decoding at test time, which adds latency, or fine tune with...
Zian Ding, Zi-Lin Zhao, Ying-Jie He et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.