Beyond Reported Accuracy: A Verifiability-Focused, Multi-Axis Review of Content and Context Based Fake News Detection
Abstract
Fake news undermines information integrity, and the proliferation of large language models (LLMs) has intensified both its generation and its detection. Prior surveys catalogue methods but rarely verify the performance figures they synthesize, restrict comparison to same-benchmark settings, or assess risk of bias, so reported accuracies are often over-read. This verifiability-focused review analyzes an enumerated, non-exhaustive corpus of 45 content- and context-based detection studies (2018-2025; 42 quantitative-core) in which each reported headline metric is traced to its primary source and confirmed against it wherever the source is accessible (36 of 45; the nine paywalled studies are flagged as reported-but-unconfirmed), and released with the corpus. Studies are grouped into five method families, Transformer/PLM (31.0%), multimodal (26.2%), LLM-based (21.4%), graph/propagation (16.7%), and traditional machine learning (4.8%), and interpreted through a four-axis framework of evidence source, detection time, data dependence, and deployment objective. Without metric conversion or pooling, reported accuracies are strongly heterogeneous and not comparable across families (Transformer 74.8-99.36%, multimodal 82.4-99.97%, graph 84.4-94.36% or 0.75-0.93 AUC, LLM-based 48.6-89.0%); within Weibo, seven multimodal models span 82.4-91.8%, and GPT-4 attains 68.2% (not the ~95% often cited) on LIAR-binary. Three findings follow. First, reported performance appears shaped as much by data and evaluation design as by architecture, so architecture-only interpretations can be misleading; near-perfect results may reflect dataset-specific separability, leakage, or favorable evaluation designs rather than superior transferable capability. Second, the field has evolved cumulatively from static content cues toward multi-source evidence integration and reasoning, so the families are complementary evidence sources rather than exclusive choices. Third, no family is optimal across scenarios, and a central open problem is sustaining reliability under drift, new languages, attacks, and AI-generated content. We derive a dependency-ordered research agenda and, from a Vietnamese case study, transferable design principles for low-resource languages.