Skip to content
Review Open access

The Detection Arms Race: A Multimodal Survey of Deepfake Forensics from Pixels to Phonemes

Sep 2026 · International Journal of Innovative Science and Research Technology · 0 citations

Abstract

Recent progress in generative deep learning—progressing from early autoencoders and GANs to modern latent diffusion models and neural vocoders—have made synthetic media generation accessible, realistic, and increasingly difficult to identify. While these tools offer productive uses in creative production and accessibility, they create serious security risks through identity fraud, automated disinformation, and unauthorized voice and video impersonation. In response, forensic detection has developed from simple pixel-level artifact checks into complex multimodal pipelines. This survey provides an organized, critical review of current deepfake detection techniques across image, video, audio, and hybrid modalities. We examine the core architectures used in detection—such as Convolutional Neural Networks, Recurrent Networks, Vision Transformers, graph-based phoneme models, and acoustic reverberation estimators—and analyze their behavior on standard benchmarks including FaceForensics++, Celeb-DF, DFDC, ASVspoof, and FakeAVCeleb. We give particular attention to cross-dataset generalization failures, practical latency constraints for on-device deployment, and the growing need for multimodal fusion to maintain detection reliability against next-generation generative models.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.