Aug 2026· International Journal of Data Mining Techniques and Applications· 0 citations· 5 references
TL;DR
A robust, explainable detection framework is presented that combines a CNN backbone for extracting spatial artifacts with an LSTM module for modeling temporal inconsistencies across frames that enhances forensic decision support and increases practical readiness for content verification systems.
Abstract
Deepfakes pose growing risks to information integrity, yet many detectors perform well only on the datasets they were trained on and remain opaque to human analysts. A robust, explainable detection framework is presented that combines a CNN backbone for extracting spatial artifacts with an LSTM module for modeling temporal inconsistencies across frames. To make decisions auditable, the architecture incorporates Grad-CAM for spatial heatmaps, SHAP for quantitative feature attribution, and LIME for local surrogate explanations. The system was trained primarily on FaceForensics++ with stratified sampling and augmentation to reduce dataset bias and evaluated on multiple external benchmarks to assess cross-domain generalization. Experimental results show strong detection metrics, such as accuracy of 96.3%, precision of 95.8%, recall of 96.7%, and an F1-score of 96.2%, along with robust performance under JPEG compression, Gaussian noise, and FGSM adversarial attacks. By coupling high detection accuracy with transparent explanations, the proposed approach enhances forensic decision support and increases practical readiness for content verification systems.
An in-depth survey of fifteen state-of-art methodologies including classical CNN models, temporal-spatial video recognition, transformer-based networks, explainable AI (XAI) models, and models that combine multimodal large language model (LLM) products are provided.
Shavnam Shavnam, Neha Dhiman· International Journal of Inn...· 0 citations
Rapid advances in deep learning technology have led to the emergence of artificial intelligence (AI) media that is very similar to reality, called deepfakes, which have the potential to pose a serious threat to information integrity and public trust. Although detection methods using Convolutional Neural Networks (CNN) have been developed, most still struggle with generalization, particularly in distinguishing modern deepfakes from non-standard original images such as selfies, which often leads to high false positive rates. This study introduces a robust detection model based on the EfficientNetB0 architecture implemented through transfer learning techniques. To improve generalization capabilities and minimize bias, we compiled a large and balanced combined dataset by combining three different public datasets (including classic deepfakes, face swaps, and many authentic selfies). The model was trained using a two-stage strategy: first for feature extraction, then refinement with a very low learning rate. The model's performance was thoroughly evaluated on stratified test data using five key metrics. The results of the experiment showed outstanding performance, achieving 99.81% accuracy and a Macro F1 score of 99.81%. Additionally, the reliability metrics ROC-AUC, Average Precision (AP), and True Positive Rate (TPR) all reached 99.99%, while the False Positive Rate (FPR) remained strictly at 1%. As proof of concept, this optimized model was implemented in a web prototype built using the Django framework, allowing users to upload images and receive classification results in real-time.
Muhammad Erico Revaldo, Burhanudin Rabbani, Wildan Humaidi et al.· Jurnal Teknoinfo· 0 citations
An ensemble-based approach, namely the X-Iv2 Ensemble approach, merging Inception ResNet v2 and Xception Net based on their complementary architectures to enhance feature extraction and classification is introduced, integrating Explainable AI (XAI) using Integrated Gradients to interpret decision-making processes.
P. P. Sudharsana, R. Rajalaxmi· Intelligent Data Analysis· 0 citations
DeepFakeBuster is presented as a confidence-calibrated adaptive ensemble for deepfake image detection by fusing together heterogeneous deep learning models built around detecting complementary forensic cues e.g., spatial inconsistencies, boundary artifacts, noise residuals, semantic consistency, and frequency-domain features.
Rachana Patil, R. Shinde, S. Patil et al.· Scientific Reports· 0 citations
This project presents an explainable deep learning framework for identifying real and AI-generated images using the NASNet architecture and achieves high detection accuracy while providing interpretable visual explanations, making it suitable for digital image verification, media authentication, and cybersecurity applications.
Panduga Mounika, Dr.CH. Buchi Reddy· American Journal of AI Cyber...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.