Skip to content
Open access

A Multi-Domain Feature Framework for Robust Deepfake Audio Detection

2026 · ITEGAM- Journal of Engineering and Technology for Industrial Applications (ITEGAM-JETIA) · 0 citations

Abstract

Audio deepfakes generated by modern text-to-speech and voice conversion systems pose serious threats to security, privacy, and trust in digital communication. This study proposes a multi-domain feature fusion framework for robust deepfake audio detection under realistic, in-the-wild conditions. A large-scale dataset comprising 31,780 audio samples, evenly split between genuine and synthetic speech and covering diverse speakers, languages, and recording environments, is utilized. Acoustic, compression-related, emotional, phase-based, prosodic, and statistical–spectral features are extracted and fused, and classification is performed using a lightweight fully connected neural network evaluated via stratified five-fold cross-validation. The proposed system achieves an average validation accuracy of 98.03% and an AUC of 0.998, demonstrating strong and stable discriminative performance. Ablation experiments and SHAP-based analysis highlight the critical role of compression-related features in enhancing robustness when combined with other feature domains. While the study focuses on audio-only detection and does not address adversarial or multimodal scenarios, the results indicate that multi-domain feature fusion offers a practical and generalizable solution for real-world deepfake audio detection, particularly in environments involving diverse codecs and synthesis techniques.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.