Skip to content

KGEMD: knowledge-guided enhanced multimodal detection for fake news

Aug 2026 · Multimedia Systems · Vol 32 · 0 citations · 45 references

TL;DR

The results indicate that explicitly modeling semantic conflict as a discriminative feature effectively improves detection precision and generalization, providing a robust solution for factual verification in complex media environments.

View source

Similar papers

Open access Jul 2026

TERN: Type-Aware Evidence Reasoning for Multimodal Fake News Detection

TERN induces latent deception types from image-side multimodal features through prototype-based clustering, uses the induced assignments as structural priors for downstream veracity prediction, disentangles type-discriminative factors from semantic content, and performs type-conditioned hierarchical reasoning over text semantics, image authenticity, and cross-modal consistency.

Mingshu Zhang, Hongyu Jin, Yuechuan Zhang et al. · 1 citation
Book Open access Aug 2026

A Serial Two-Stage Framework for Robust Multimodal Fake News Detection via Adaptive Reasoning

DAR-Lite is proposed, a serial two-stage framework that rethinks the detection pipeline through explicit decoupling of representation denoising and contextual reasoning, and achieves a favorable balance between detection performance and computational cost.

Maolin Wang, Ziting Mai, Zichun Liu et al. · 0 citations
Open access Jul 2026

SMD-Net: Selective and Multiway Differential Perception-Enhanced Network for Early Multimodal Rumor Detection

The rapid dissemination of rumors on social media at their early stages poses significant threats to public safety and social stability. While response-based methods usually depend on user comments and reposts and therefore suffer from inherent latency, existing content-based methods still struggle to extract discriminative evidence from noisy short texts and subtle visual inconsistencies under zero-response conditions. To address this issue, we propose SMD-Net, a multimodal framework for early zero-response rumor detection. In the textual branch, a selective state-space encoder is used to model fragmented and noisy posts. In the visual branch, an enhanced TransXNet backbone is designed to improve the representation of fine-grained suspicious patterns and cross-layer feature interactions. An adaptive gated fusion module is further introduced to integrate textual and visual features for final prediction. Experiments on the Weibo and PHEME datasets show that SMD-Net outperforms the compared content-based baselines, achieving 92.60% accuracy on Weibo and 90.27% accuracy on PHEME under the strict zero-response setting. These results suggest that the proposed framework provides an effective solution for early multimodal rumor detection when propagation-based evidence is unavailable.

Zheng-Nan Qiao, Zhe-Kang Yang, Xianguo Zhang · 0 citations
Open access Aug 2026

Context-aware multimodal reasoning for explainable Bengali fake news detection using vision-language models

In particular, the rapid dissemination of fake news using social media platforms has posed a severe threat to social stability and information reliability in language groups that lack sufficient resources. In addition, it has become difficult to automatically identify fake news in the case of Bengali digital media because fake news is mostly distributed using multimodal data like memes, screenshots, posters, photos with texts, etc. The effectiveness of the existing fake news detection methods is somewhat hindered by their inability to provide explainability and their focus mainly on either textual or visual data. This study proposes a context-aware multimodal reasoning approach for explainable Bengali fake news detection. The proposed model incorporates EasyOCR for Bengali textual information extraction, ResNet-50 and ViT for supplementing visual feature learning, and Qwen2-VL-2B-Instruct for multimodal semantic inference. The proposed approach can detect semantic contradictions and associations between textual assertions and image information through matching textual and visual data based on a context-sensitive fusion technique. This methodology generates interpretable explanations beyond the fake/real binary categorization to enhance user trust in automated outcomes. According to an experiment conducted using a multimodal Bengali fake news dataset, the proposed approach surpasses conventional CNN, transformer, and unimodal baselines on several performance metrics. Results indicate how context-based multimodal reasoning can improve the efficiency and robustness of the model, along with making it more interpretable. The proposed method serves as an encouraging route towards the identification of fake news in multiple low resource languages, as well as in combating misinformation in the Bengali digital ecosystem.

Shillpi Mishrra, Sauvik Bal, Arpita Dhar et al. · 0 citations
Book Open access Aug 2026

MAR: Metacognitive Agentic Reasoning for Multimodal Fake News Detection

Multimodal fake news combining text and images has become increasingly prevalent, fueled by the rapid dissemination on social media. Existing approaches predominantly rely on supervised learning–driven small multimodal language models, yet they are constrained by the knowledge scope and logical reasoning capabilities limited by training data: their performance degrades significantly when encountering novel content not covered in the training data, and they typically provide only uninterpretable binary true/false predictions. To address these challenges, we propose the Metacognitive Agentic Reasoning for Multimodal Fake News Detection(MAR), a framework that integrates multi-agent metacognitive debate with external knowledge retrieval to enhance the accuracy, generalization, and interpretability of multimodal fake news detection. At its core, MAR introduces a news-domain informed and metacognition-inspired multi-agent reasoning mechanism: it first generates several prior pseudo-labels of the news domain, and defines the characteristics of two opposing agents (i.e., Believer and Skeptic) leveraging these pseudo-labels; then, the Believer and Skeptic go through a three-stage ''Draft – Self-Critique – Refine'' interactive debate simulating humans' learning behavior, called the metacognitive debate. Specifically, MAR first generates an initial evidence-grounded judgment (Draft), then critically reflects on their own and others' arguments (Self-Critique), and finally iteratively refines their conclusions by incorporating internal or external evidence (Refine). To improve factual reliability, the framework dynamically retrieves trustworthy external knowledge via web and reverse image searches, thereby mitigating the hallucinations inherent in large language models. Experiments show that MAR achieves state-of-the-art performance on two benchmarks and significantly outperforms existing methods in terms of generalization and interpretability. The source code is available at https://github.com/Averdgr/MAR_KDD.

Wenyu Chen, Hengbing Dong, Junhao Wa et al. · 0 citations
Open access Aug 2026

A Multimodal Fake News Detection Model Based on Adaptive Binary Osprey Optimization Algorithm and Cross-Modal Disentangled Fusion

An Adaptive Binary Osprey Optimization Algorithm and Cross-modal Disentangled Fusion model (ABOOA-CDF) is proposed that effectively improves detection performance and effectiveness in feature optimization, cross-modal relation modeling, and semantic fusion.

Xuran Dai, Guozan Lu, Jiaxue Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.