Skip to content

CRAF: Cross-View Residual-Aware Fusion for Deepfake Speech Detection

Sep 2026 · 0 citations · 23 references
Computer Science Engineering

TL;DR

CRAF, a cross-view residual-aware fusion framework that uses ALLM-guided cross-view attention to enrich SSL representations and adopts ALLM as a high-level reference to separate ALLM-explainable information from complementary SSL residual information, demonstrates robustness to unseen spoofing attacks.

Abstract

Recent advances in speech synthesis and voice conversion have made deepfake speech increasingly realistic, making generalization to unseen spoofing attacks a critical challenge. Pretrained speech and audio models offer a promising direction for improving robustness to such unseen attacks. Self-supervised learning (SSL) models capture fine-grained, low-level acoustic characteristics, whereas Auditory Large Language Models (ALLMs) provide higher-level contextual representations. These complementary views can provide useful cues for improving generalization to unseen attacks. However, direct fusion does not explicitly disentangle information shared across the two views from view-specific complementary information, limiting effective cross-view integration. To address this, we propose CRAF, a cross-view residual-aware fusion framework that uses ALLM-guided cross-view attention to enrich SSL representations and adopts ALLM as a high-level reference to separate ALLM-explainable information from complementary SSL residual information. The residual is selectively refined through adaptive gating and integrated through SSL-primary fusion. Experiments on ASVspoof 5 show that CRAF with Kimi-Audio achieves an EER of 5.96% and a minDCF of 0.1192, demonstrating robustness to unseen spoofing attacks.

View source

Similar papers

Preprint Sep 2026

Disentangled Global-Local Feature Learning with E-Branchformer for Audio Deepfake Detection

The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervis...

Phuong Dat, Học Thủ, T. Nguyễn et al. · 0 citations
#artificial intelligence Preprint Sep 2026

GLAD: Global-Local Adaptive Detector for Robust Speech Deepfake Detection

This paper proposes the Global-Local Adaptive Detector (GLAD), a Hierarchical Global-Local backbone that explicitly bridges the granularity gap by fusing global linguistic and acoustic features with fine-grained local signal details and introduces SaniBoost, a composite data augmentation strategy for robust signal stan...

Ze-Lin Zhao, Guan-Jie Huang, D. H. K. Tsang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection

The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation mechanism changes....

Jia-Jun Xu, Meng-Lu Li, Xiao-Ping Zhang · 0 citations
Preprint Sep 2026

Domain-Adaptive Dual-Gating Mixture of Experts for Generalizable Speech Deepfake Detection

Recent advances in speech deepfake detection (SDD) have leveraged the Mixture of Experts (MoE) to enhance generalization capacity. However, existing gating networks often overlook the acoustic and temporal cues of deepfakes. In this work, we propose a novel domain-adaptive dual-gating MoE (DADGMoE) framework for SDD un...

Si-Qing Qin, Zhe Li, K. Lee et al. · 0 citations
Preprint Aug 2026

Beyond Speech: Dual-Domain SSL Fusion for Unified All-Type Audio Deepfake Detection

Unified all-type audio deepfake detection aims to determine whether an input clip is real or fake when its audio type may be speech, environmental sound, singing voice, or music. Existing speech-centric or type-dependent solutions are insufficient for this setting because the test-time audio type is unknown, while the...

Cunhang Fan, Jun-Qin Cao, Tian Gao et al. · 0 citations
Open access Aug 2026

EMBNet: Multi-scale feature learning with efficient channel attention for deepfake speech detection.

EMBNet is proposed, a task-oriented deepfake speech detection framework that integrates efficient channel attention (ECA) with a multi-scale bottleneck to enhance the representation of subtle acoustic anomalies and may support forensic audio authenticity assessment by improving the discrimination between genuine and ma...

Haitao Yang, Fen Li, Xin Cai et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.