Benchmarking Deepfake Detectors: Comparative Analysis of Spatial, Temporal, and Hybrid Architectures
Generative AI advancements are rapidly producing highly realistic deepfake videos, which raises major concerns for digital authenticity and security. This paper presents a comparative analysis of representative spatial, temporal, and hybrid deepfake detection architectures. Selected representative models include a lightweight MobileNetV2 + TinyViT hybrid approach; 2D Xception-based spatial model; 3D temporal video model; and the Knowledge-Guided Temporal Transformer (KGTT). Experimental results suggest that hybrid architectures provide the most balanced trade-off among the evaluated models, providing the best compromise between accuracy and computational efficiency. The KGTT model achieved frame-level accuracy exceeding 95%, while the MobileNetV2 + TinyViT hybrid achieved 80.46% accuracy with improved computational efficiency. Further, purely temporal models depend on large-scale, balanced datasets. Collectively, the findings show that integrating complementary feature domains is critical for developing robust and generalizable deepfake detection methods. This study provides a greater under-standing of the strengths and weaknesses of current approaches to deepfake detection, along with insights into promising areas for future research.