Skip to content
Review Open access

Advancements of Audio Unimodal Deep Faking Detection Technology

Aug 2026 · Mathematical Modeling and Algorithm Application · 0 citations · 23 references

TL;DR

This paper systematically reviews the types of audio forgery, detection principles, authoritative datasets and evaluation indicators, compares and analyzes traditional detection methods with deep learning detection techniques, and points out the core challenges in generalization, robustness, etc. of current methods.

Abstract

In recent years, the generative artificial intelligence technology has undergone rapid iterations, leading to the widespread abuse of audio deep forgery methods such as speech synthesis and speech conversion, which seriously threaten information security and social credibility. Among various countermeasures, audio single-modal deep forgery detection has advantages such as lightweight and strong real-time performance, and is a key technology for security protection in pure audio scenarios, with significant research and application value. This paper systematically reviews the types of audio forgery, detection principles, authoritative datasets and evaluation indicators, compares and analyzes traditional detection methods with deep learning detection techniques, focuses on elaborating the characteristics and applicable scenarios of four mainstream detection models, and points out the core challenges in generalization, robustness, etc. of current methods. Finally, it looks forward to the development trends of this field towards generalization, high robustness, lightweight and forgery traceability, which can provide references for related research.

Read PDF

Similar papers

Conference Jul 2026

Deepfake tampering detection based on multimodal large model

As deep learning generation technologies continue to evolve, Deepfake technology has become a major threat to information security, posing significant challenges to tampering detection. Existing detection methods for deep forgeries generally suffer from insufficient generalization capability and interpretability. To address these issues, this paper proposes a Deepfake tampering detection method based on a multimodal large model, termed LLM-FNP. This method integrates frequency-domain noise perception with visual-language understanding. First, the BayarConv module is employed to extract frequency-domain noise features from images, capturing anomalous noise patterns left by the forgery process, and a dual-branch encoder is used to process the original RGB images and frequency-domain feature maps separately. Second, a cross-attention mechanism is applied to achieve cross-modal fusion. Finally, the fused features are processed by a fine-tuned LLaVA large model, which outputs the detection result along with interpretable text. Experimental results on the DFDCP dataset show that the proposed method achieves an accuracy of 78.54% and an AUC of 81.50%. To evaluate the generalization capability of the method, cross-dataset testing on Celeb-DF and WildDeepfake yields accuracies of 73.5% and 71.8%, and AUCs of 76.4% and 75.1%, respectively, all outperforming classical methods such as ResNet+LSTM, further validating its strong generalization ability.

Shuyi Tao, Jiaye Li, Cuiling Jiang · 0 citations
Open access Aug 2026

Deepguardnet: A Resnet-Based Hybrid Framework for Intelligent Deepfake Image and Video Authentication

The evolution of sophisticated generative artificial intelligence has led to the rapid development of very realistic manipulated images and videos, posing substantial risks for digital trust, cyber security, and multimedia authenticity. Advanced Deepfake generation technologies result in the creation of believable forgery media which become hard to differentiate from authentic media; this leads to misinformation, identity spoofing, and digital scams. Therefore, precise and effective authentication of multimedia becomes an imperative requirement for digital forensics investigation and online content authentication. This research paper presents a ResNet-powered deep feature learning approach for detecting Deepfake images and videos. The suggested approach normalizes and resizes images, while videos are decomposed into frames for thorough spatial-temporal analysis. The hybrid convolutional neural network model, which is built on top of ResNet architecture, extracts discriminative features that represent subtle manipulation traces, face texture inconsistency, and structural abnormalities. In addition, inverted residual blocks and linear bottlenecks are used to increase computational efficiency. The deep learning-based feature extraction process is then followed by the classification stage to distinguish between genuine multimedia content and forged multimedia content. From experimental studies, it can be shown that the proposed framework helps to enhance the detection rate, robustness toward new Deepfake methods, and enables real-time implementation. The research provides an effective solution for multimedia authentication applications.

Bella Inba Suganthi V, S. Jose · 0 citations
Open access Aug 2026

Multimodal Deepfake Detection for Digital Forensics: A Robust Audio-Visual Inconsistency Approach for Evidence Integrity

This research provides a resilient forensic layer for digital identity verification, ensuring evidence integrity in the GenAI era by proposing a novel Multimodal Deep Learning framework designed to detect high-fidelity Deepfakes by exploiting audio visual temporal inconsistencies.

H. Truong, Dung The Luong, Tuan Tran · 0 citations
Review Open access Jul 2026

Deep Learning-Based Fake Image Detection Using Transfer Learning: A Systematic Review

This review presents a comprehensive analysis of recent deep learning and transfer learning techniques for fake image detection, examining widely adopted convolutional neural network architectures, benchmark datasets, evaluation metrics, and current research developments.

Nisha Parveen, Anjali Saxena · 0 citations
Review Open access Jul 2026

Multimodal transformer-based watermarking for deepfake detection and digital media authentication: current progress, challenges, and future directions

The rapid advancement of deepfake generation technologies has fundamentally outpaced the forensic tools designed to detect and authenticate digital media. Traditional watermarking methods, while foundational, were not conceived for the adversarial complexity of multimodal content ecosystems where video, audio, and image signals are increasingly synthesized, blended, and redistributed at scale. This gap has made reliable media provenance one of the most pressing open problems in applied artificial intelligence. This mini review surveys the current landscape of transformer-based approaches to digital watermarking and deepfake detection, with a focus on their capacity to operate across multiple modalities within unified architectures. We trace the progression from classical signal-based watermarking to attention-driven deep learning frameworks, highlighting where transformer models offer meaningful resilience gains over legacy methods. We further examine emerging efforts to consolidate watermark embedding, forgery detection, and content authentication into integrated pipelines and discuss why such unification is both technically advantageous and practically necessary. The review closes by mapping the field's most consequential open challenges including cross-modal generalization, adversarial robustness, and benchmark scarcity and identifying the directions most likely to yield progress.

Kok Swee Sim, M. Islam · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.