Skip to content
Open access

Multimodal Deepfake Detection for Digital Forensics: A Robust Audio-Visual Inconsistency Approach for Evidence Integrity

Aug 2026 · Journal of Science and Technology on Information security · pp. 5-26 · 0 citations · 35 references

TL;DR

This research provides a resilient forensic layer for digital identity verification, ensuring evidence integrity in the GenAI era by proposing a novel Multimodal Deep Learning framework designed to detect high-fidelity Deepfakes by exploiting audio visual temporal inconsistencies.

Abstract

The rapid proliferation of Generative AI (GenAI) has democratized the creation of hyper-realistic multimedia forgeries, posing severe threats to electronic Know Your Customer (eKYC) systems and digital forensic investigations. While visual synthesis has reached near-perfection, maintaining precise synchronization between lip movements (visemes) and speech signals (phonemes) remains a formidable challenge. To address this, we propose a novel Multimodal Deep Learning framework designed to detect high-fidelity Deepfakes by exploiting audio visual temporal inconsistencies. Beyond traditional feature fusion, our architecture integrates a Contrastive Synchronization Loss with a Transformer based Cross-Modal Attention mechanism. This hybrid objective explicitly enforces intra-class compactness for authentic pairs while amplifying the distance for asynchronous forgeries. Extensive experiments on FaceForensics++, DFDC, and a custom Vietnamese dataset (Vn-eKYC-Aug) demonstrate that our model achieves state-of-the-art performance, maintaining high robustness against video compression and environmental noise, though operational efficacy remains sensitive to extreme low-light conditions and diverse regional dialects. This research provides a resilient forensic layer for digital identity verification, ensuring evidence integrity in the GenAI era.

Read PDF

Similar papers

Preprint Aug 2026

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video. Yet existing research either focuses on visual manipulation, addresses speech detection in isolation, or conflates speech and non-speech audio as a single undifferentiated audio stream, overlooking the distinct forensic challenges posed by background audio. This conflation is consequential: the two acoustic components arise from fundamentally different generative mechanisms, exhibit distinct artifact profiles, and pose different challenges to detection systems. We introduce MADBench, the first benchmark that treats speech and environmental audio as distinct acoustic components, enabling component-aware evaluation of audio deepfake detection across independently manipulated forgery sources. We benchmark representative state-of-the-art detectors and multimodal large language models under a unified protocol. Our experiments reveal that environmental audio manipulation is more detectable than synthetic speech across general-purpose encoders, while existing pretrained detectors fail on both acoustic components, and manipulated environmental audio asymmetrically degrades speech deepfake detection, findings entirely invisible under the single-label paradigm of prior benchmarks. MADBench establishes a rigorous foundation for future research into robust, component-aware audio deepfake detection.

Yanqiu Li, Yang Xiao, Jisheng Bai et al. · 0 citations
Open access Aug 2026

Deepguardnet: A Resnet-Based Hybrid Framework for Intelligent Deepfake Image and Video Authentication

The evolution of sophisticated generative artificial intelligence has led to the rapid development of very realistic manipulated images and videos, posing substantial risks for digital trust, cyber security, and multimedia authenticity. Advanced Deepfake generation technologies result in the creation of believable forgery media which become hard to differentiate from authentic media; this leads to misinformation, identity spoofing, and digital scams. Therefore, precise and effective authentication of multimedia becomes an imperative requirement for digital forensics investigation and online content authentication. This research paper presents a ResNet-powered deep feature learning approach for detecting Deepfake images and videos. The suggested approach normalizes and resizes images, while videos are decomposed into frames for thorough spatial-temporal analysis. The hybrid convolutional neural network model, which is built on top of ResNet architecture, extracts discriminative features that represent subtle manipulation traces, face texture inconsistency, and structural abnormalities. In addition, inverted residual blocks and linear bottlenecks are used to increase computational efficiency. The deep learning-based feature extraction process is then followed by the classification stage to distinguish between genuine multimedia content and forged multimedia content. From experimental studies, it can be shown that the proposed framework helps to enhance the detection rate, robustness toward new Deepfake methods, and enables real-time implementation. The research provides an effective solution for multimedia authentication applications.

Bella Inba Suganthi V, S. Jose · 0 citations
Open access 2026

Multimodal Fusion and Explainable Deep Learning for Synthetic Voice and Vishing Detection

Synthetic speech generation and voice-cloning technologies have achieved unprecedented levels of realism, enabling numerous applications in accessibility, virtual assistants, and media production. However, these advancements also introduce significant risks, including identity fraud, impersonation attacks, misinformation, and security breaches. This paper proposes a multimodal fusion framework for synthetic voice detection that combines handcrafted acoustic features with deep spectrogram representations to improve detection robustness and generalization. The proposed architecture employs a Convolutional Neural Network–Bidirectional Long ShortTerm Memory (CNN-BiLSTM) network to capture both spectral artifacts and temporal inconsistencies characteristic of AI-generated speech. To enhance transparency and interpretability, an explainability module incorporating attention visualization and feature attribution techniques is integrated into the detection pipeline. Furthermore, the framework is deployed through a real-time inference interface, demonstrating its practical applicability in cybersecurity, digital forensics, and media authentication scenarios. The findings highlight the effectiveness of combining deep learning, multimodal feature fusion, and explainable artificial intelligence to address the growing challenge of synthetic speech detection.

Mahima Bg, Pallavi Gb · 0 citations
Jul 2026

Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection

A detection framework is proposed that extracts per-video rPPG wave- forms via RhythmFormer and trains a suite of lightweight classifiers to distinguish real from synthesized physiologi- cal signals and shows that detec- tion difficulty is strongly method-dependent.

Othmane Harraq, Tamer Aldwairi · 0 citations
Review Open access Aug 2026

Advancements of Audio Unimodal Deep Faking Detection Technology

This paper systematically reviews the types of audio forgery, detection principles, authoritative datasets and evaluation indicators, compares and analyzes traditional detection methods with deep learning detection techniques, and points out the core challenges in generalization, robustness, etc. of current methods.

Minglei Zhu · 0 citations
Preprint Aug 2026

Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow

A generative detector that localizes tampering by estimating the local restoration cost required to align a query image with authentic visual-text statistics, rather than by learning forgery-specific decision boundaries is proposed, and Sparse-Constraint Rectified Flow is introduced, a detector-oriented adaptation of Flow Matching for spatially sparse anomaly localization.

Jiangling Zhang, Shuxuan Gao, Zeyu Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.