Aug 2026· Journal of King Saud University: Computer and Information Sciences· Vol 38· 0 citations· 45 references
TL;DR
LCFND, a staged and fully specified bounded-evidence verification pipeline for localized conflict discovery, structured evidence generation, and graph-based verification, is presented, better suited to offline or high-risk verification than to large-scale real-time screening.
Abstract
Multimodal fake news is difficult to detect when text and image remain topically aligned but diverge in local factual details such as entities, locations, time, or event attributes. Existing methods often rely on global cross-modal fusion and may therefore overlook these localized contradictions. We present LCFND, a staged and fully specified bounded-evidence verification pipeline for localized conflict discovery, structured evidence generation, and graph-based verification. Optimal transport identifies high-conflict token-region pairs and constructs localized evidence packets. A parameter-efficient LLM trained with OT-grounded structured targets and training-split label-informed weak relation calibration then converts these packets into structured evidence, including claims, visual evidence, contradiction types, and grounded rationales. A heterogeneous evidence graph integrates OT conflict priors and LLM-generated evidence for veracity prediction, and all graph-construction rules, training objectives, and evaluation protocols are stated explicitly for reproducibility. Under a controlled matched-backbone protocol on Weibo, Twitter, and GossipCop, LCFND improves F1 over the top matched-backbone baseline MGCA by 1.47, 1.73, and 1.67 points, respectively. The gain remains positive when the visual backbone is strengthened to ViT-B/16 or Swin-T. Manual evidence evaluation on 500 samples per dataset reports 85.9–87.3% claim grounding and 78.1–80.2% contradiction-type correctness. Under the P3b matched-input protocol, LCFND remains above FND-LLM-M and MMRGV-lite, and improves average cross-dataset transfer F1 from 67.10 to 69.98. Full online inference takes 1.47 s per sample in our setting, so the framework is better suited to offline or high-risk verification than to large-scale real-time screening.
CAER introduces a span-grounded evidence router that transforms claim representations into soft textual queries and retrieves corresponding evidence from frozen visual tokens, enabling fine-grained conflict estimation and design a dual-prefix expert routing mechanism that learns separate experts for visually supported and contradicted inputs, enabling conflict-aware generation through explicit expert selection.
Zi-Xuan Liu, Juntao Cai, Xiaoxu Cai et al.· arXiv.org· 0 citations
The results indicate that explicitly modeling semantic conflict as a discriminative feature effectively improves detection precision and generalization, providing a robust solution for factual verification in complex media environments.
Zi-Heng Wang, Junfang Song, Shuyu Wang et al.· Multimedia Systems· 0 citations
Video misinformation detection is often approached through global multimodal fusion or free-form multimodal reasoning. Both paradigms can under-represent localized authenticity cues that arise from coupled interactions among query phrases, contextual text, and short temporal spans of frames. Because such interactions are inherently higher-order, pairwise graph formulations are insufficient to capture multi-way cross-modal dependencies, whereas hypergraphs offer a suitable representation for these relations. We propose HyperClaim, a discriminative temporal hypergraph framework for sample-level authenticity classification. Using the title or benchmark-provided paired text as a claim-like query, HyperClaim constructs a sparse heterogeneous hypergraph over query tokens, evidence tokens, and sampled frames; applies confidence-aware filtering and source budgeting to form compact text-frame and short-range temporal evidence units; performs adaptive soft-incidence reasoning with residual text-video calibration; and aggregates textual, visual, and hyperedge states through a discrepancy-aware readout. Without relying on generated rationales or external tool calls, HyperClaim preserves fine-grained cross-modal and temporal structure that global fusion tends to flatten. Under the FactGuard temporal protocol, it achieves 83.7%, 82.0%, and 87.3% accuracy on FakeSV, FakeTT, and FakeVV, respectively, outperforming strong discriminative and reasoning-centric baselines. Learned incidence and attention weights further reveal token- and frame-level structure.
Xiangbo Wang, Jiasheng Zhang, Xingtong Yu et al.· 0 citations
Fake news increasingly relies on cross-modal image-text forgeries, making transparent and verifiable reasoning chains an urgent need for Detecting and Grounding Multi-Modal Media Manipulation (DGM4). Existing methods produce black-box detection results without any decision rationale, limiting their reliability in forensic practice. Multi-modal Large Language Models (MLLMs) offer a natural path toward explainability, but applying them to DGM4 raises two difficulties. First, models tend to generate explanations disconnected from predicted evidence locations, producing unverified attribution. Second, enforcing evidence-conclusion consistency requires active optimization, yet uniform training signals fail to distinguish localization tokens from classification tokens, making multi-head joint training unreliable. We propose a multi-modal manipulation detector based on an Evidence-Grounded Forensic Reasoning (EFR) framework. EFR introduces an Anchor-and-Verify reasoning chain that enforces modality-isolated perception before cross-modal comparison, with conclusion coordinates as explicit anchors to which downstream evidence must spatially correspond. A verifiable reward system then enforces evidence-conclusion consistency during training, while a Modality-Decoupled Advantage (MDA) routing mechanism mitigats credit misassignment across prediction tasks. Experiments show that EFR achieves state-of-the-art performance while producing structured forensic reasoning records that explicitly bind explanations to evidence.
Multimodal Large Language Models (MLLMs) have achieved strong performance on OCR-centric document understanding and text-rich visual reasoning benchmarks. Yet existing evaluations largely assume that every task is valid and answerable. In real-world OCR scenarios, this assumption often fails: questions may rely on illegible text, occluded evidence, nonexistent visual targets, contradictory premises, or missing variables. We study this reliability gap as OCR-grounded Task Verification: before answering, a model should determine whether the Image Premise (IP), Textual Premise (TP), and Question (Q) jointly define an executable task. We introduce VeriOCRBench, a 1,800-sample human-verified benchmark built from source images drawn from 8 OCR-related datasets and spanning 8 real-world image domains, with controlled, image-grounded diagnostic tasks. It contains 1,600 trap-injected invalid tasks across 8 trap types and four verification dimensions---Visual, Contextual, Factual, and Logical---plus 200 trap-free controls for measuring over-refusal. Built with a Visual Atomic Fact (VAF)-anchored pipeline and full human auditing, VeriOCRBench enables decoupled evaluation of task verification, root-cause diagnosis, and over-refusal. Evaluating 15 leading MLLMs reveals persistent blind compliance, diagnosis failures, and prompt-induced over-refusal, exposing a critical reliability gap in current OCR reasoning systems. The code is available at: https://github.com/zy001122/Beyond-Blind-Compliance.
Experiments with representative closed-source and open-source MLLMs show that OCR-grounded meta-reasoning remains far from saturated: models struggle with visible-rule application and layout-sensitive inference, while process-compliant rationales can accompany incorrect final answers under exact-match evaluation.
Geng-Xu Li, Yuan Wu, Yi Chang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.