Skip to content

Discourse Structure as an Interpretable Signal for Detecting Hallucinated Chain-of-Thought Reasoning in Large Language Models

· 0 citations · 39 references

TL;DR

This paper studies hallucinated CoT as a discourse-structural phenomenon, not only a factual one, and suggests that discourse structure provides an interpretable signal for detecting reasoning hallucinations and complements existing factuality and uncertainty-based hallucination detectors.

View source

Similar papers

Book Open access Aug 2026

Who's Adam? Benchmarking Hallucinations in Scientific Dialogue

ADAM-Bench (Auditing Dialogue Assertions with Multimodal Evidence), a benchmark for paper-grounded hallucinations in scientific dialogue, is introduced and two tasks are defined: hallucination detection and minimal evidence set localization.

Zexing Zhang, Tianyang Lei, Kewei Yang et al. · 0 citations
Open access Aug 2026

Layer-wise symbolic attention instability as a diagnostic signal for hallucination in large language models

A unified symbolic, behavioral, and mechanistic framework that connects symbolic triggers with internal failure dynamics in transformer architectures and provides an interpretable basis for diagnosing and stabilizing symbolic reasoning in LLMs is introduced.

Naveen Lamba, Sanju Tiwari, Manas Gaur · 0 citations
Preprint Aug 2026

Hallucination Span Detection with Input-Side Evidence Alignment

This work introduces the task of hallucination span detection with input-side evidence alignment, which jointly identifies hallucinated spans and aligns output tokens with the corresponding input evidence.

Miyu Yamada, Yuki Arase · 0 citations
#artificial intelligence Preprint Aug 2026

EviAnchor: Mitigating Hallucinations in Large Vision-Language Models via Regional Visual Evidence Compensation

Large vision-language models (LVLMs) frequently generate content unsupported by visual inputs. Preliminary experiments show that visual evidence is primarily incorporated into answer-side representations in early-to-middle decoder layers, while its direct influence progressively weakens in later layers. This attenuation suggests that visual evidence acquired earlier may be insufficiently utilized during subsequent generation. Based on this observation, we propose EviAnchor, a training-free and single-branch inference framework that preserves and reactivates visual evidence throughout generation. EviAnchor introduces Regional Evidence Anchor (REA) slots to progressively aggregate dense visual tokens into spatially structured representations. It then strengthens the current decision state's access to these visual anchors through decision-conditioned evidence routing, mitigating excessive dependence on textual context. Finally, the model resumes its native Transformer computation to integrate the retrieved visual evidence with question semantics and generation history. Experiments across POPE, CHAIR, and MMHal-Bench demonstrate consistent improvements in visual grounding.

Sihang Jia, Shuliang Liu, Song-Bo Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Leveraging Low-Level Symbolic Competences for Unsupervised Grounding in Hallucination Detection

Hallucination-where a language model generates outputs that are factually incorrect or unsupported by the source-is a major challenge for both prompted and fine-tuned language models. Detecting hallucinations is difficult due to the opaque reasoning processes of LLMs, which often provide little insight into why a model's output may be inaccurate. In this work, we investigate whether an LLM can use an alternative, low level, symbolic competence such as SQL for unsupervised hallucination detection in some high level task. For this, we make an LLM build an SQL database from reference documents. This SQL database is then used for reasoning over the reference and the sampled response in a hallucination detection pipeline that is grounded in the database, thereby providing a neurosymbolic checkup. On RAGTruth and DiaHalu hallucination detection datasets, we find that our approach improves on direct prediction and competes with state-of-the-art hallucination detection methods, while not requiring domain-specific fine-tuning. Instead it relies on a low-level general competence already present in LLMs. This warrants further investigation of low-level LLM competences in neurosymbolic approaches.

Renato Vukovic, Hsien-Chin Lin, Carel van Niekerk et al. · 0 citations
Open access Jul 2026

AI Hallucination as Epistemic Proliferation

The Epistemic Proliferation Model is conceptualised as an ethical failure of restraint, that is, the tendency of fluent models to keep producing confident, expert-sounding language after evidential grounding has become weak or unverifiable.

Indrajith P. Karunanayaka · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.