Skip to content
Open access

Detection of Face-Swap Based Deepfake Videos Using Hybrid CNN-LSTM Architecture

Unknown authors
Aug 2026 · International Journal for Research in Applied Science and Engineering Technology · 0 citations

Abstract

The development of deepfake technologies due to breakthroughs in AI and deep learning allows producing highly realistic manipulated videos and audio, thus posing a threat to misinformation and digital security. Despite deepfake technology having several legitimate uses, including use in the media industry, its inappropriate use for distributing fake news, impersonation, and cyber attacks demands the development of reliable detection techniques. The current state-of-the-art approaches of detecting deepfakes mostly utilize spatial or temporal analysis based on CNNs and RNNs; however, most existing approaches fail to generalize well and are unable to recognize more complicated manipulations on a variety of different data sets. This paper aims at developing a novel multimodal approach to deepfake detection, integrating spatial, temporal, and audio features. The proposed system makes use of CNN-based architectures to extract spatial information from the input images, transformers to capture the temporal information, and Mel-frequency cepstral coefficients (MFCC) to analyze the audio data. These heterogeneous features are combined using attention-based learning to improve the classification accuracy. The proposed method was tested on various benchmarking datasets, including FaceForensics++, DeepFake Detection Challenge (DFDC), and Celeb-DF, yielding higher accuracy than current methods. Experimental results indicate that the combination of multimodal features enhances the detection capacity. The proposed model is highly efficient in addressing deepfake challenges in the digital forensic field.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.