EHR (Emotion Hallucination Rate), an evaluator that quantifies emotion hallucinations across six facets, and HMER (Hallucination-aware Memory-guided Emotion Reasoning), a training-free framework for emotion hallucination mitigation, which enables fine-grained mitigation across diverse hallucination facets.
Abstract
Multimodal large language models (MLLMs) have shown strong potential in open-ended emotion understanding, yet they often generate emotion hallucinations. Evaluating such hallucinations is particularly challenging for two reasons. First, emotion understanding spans multiple cognitive facets, from multimodal perception to psychological reasoning. Second, emotional interpretations are expressed in free-form language, making existing closed-ended protocols insufficient for evaluation. To address these challenges, we introduce EHR (Emotion Hallucination Rate), an evaluator that quantifies emotion hallucinations across six facets: expression, action, audio, instinct, logic, and conclusion. Using EHR, we reveal that existing mitigation methods often reduce hallucinations in some facets while aggravating them in others, exposing the limitation of coarse-grained correction and the need for facet-aware localization and mitigation. Motivated by this finding, we propose HMER (Hallucination-aware Memory-guided Emotion Reasoning), a training-free framework for emotion hallucination mitigation. HMER maintains a Hallucination Memory that records localized hallucinated claims and enables targeted logit rectification, together with an Anchor Memory that preserves reliable intermediate reasoning states to stabilize subsequent generation. By selectively suppressing unreliable cues while preserving trustworthy reasoning context, HMER enables fine-grained mitigation across diverse hallucination facets. Extensive experiments on 19 MLLMs demonstrate the prevalence of emotion hallucinations and the effectiveness of our framework across diverse model architectures.
Large language models (LLMs) have achieved significant advancements in natural language processing tasks, but they remain prone to generating hallucinations—outputs that are logically inconsistent or factually incorrect. While previous research has primarily focused on hallucinations in affirmative contexts, how negate...
Jaehyung Seo, Hyeonseok Moon, Heu-Jeoung Lim· ACM Transactions on Knowledg...· 0 citations
Large language models are increasingly used for financial question answering, while they are prone to generating hallucinated content. In this research, we propose a multi-signal framework for hallucination detection and mitigation. Our framework combines six signals (entailment, semantic similarity, claim verification...
Hallucinations in large language models (LLMs) are always seen as limitations. However, could they also be a source of creativity? This survey explores this possibility, suggesting that hallucinations may contribute to LLM application by fostering creativity. Hallucinations are not treated as creativity perse; rather,...
Xuhui Jiang, Yi Liu, Ying-Han Shen et al.· AI Plus· 0 citations
This work introduces MISHAP-Bench, a comprehensive benchmark with 12,000 challenging open-ended question-audio pairs and a rigorous evaluation pipeline covering two hallucination categories, and proposes a groundedness judge that uses reference rubrics and judge prompts guided by human annotations.
Wen-Soi Zhi, Giulio Segalini, Jian-Jia Chen et al.· 0 citations
It is argued that hallucination mitigation should be evaluated as a faithfulness--informativeness--capability trade-off rather than through hallucination scores alone, because improvements on hallucination benchmarks do not reliably transfer to broader multimodal capabilities.
Mehrdad Fazli, Sina Mansouri, Mohit Marvania et al.· 0 citations
Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and temperature-scaled calibration for response-level ha...
Varun Teja Chundru, Debasmita Biswas· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.