Aug 2026· Neural Networks· Vol 205 Pt B, pp.
109431
· 0 citations· 34 references
Medicine
TL;DR
DSF-Net is a novel neural network that introduces a dual-strategy fusion approach for the AVSELD task, designed for robust and computationally efficient multi-modal comprehension.
Abstract
Audio-visual sound event localization and detection (AVSELD) seeks to identify and locate sound-emitting objects by leveraging both audio and visual data. Current methods primarily rely on convolutional neural networks (CNNs), whose constrained receptive fields limit their ability to capture broader contextual information. Although Transformer-based architectures exhibit considerable proficiency in capturing global contextual information, their efficacy is impeded by the quadratic computational complexity associated with processing long-range dependencies. This poses a significant bottleneck, particularly in scenarios involving longer sequence lengths. To overcome this limitation, we propose DSF-Net, a novel neural network that introduces a dual-strategy fusion approach for the AVSELD task. Built upon an efficient state-space model backbone to ensure linear complexity, DSF-Net is designed for robust and computationally efficient multi-modal comprehension. The proposed dual strategies consist of: (1) an Adaptive Frequency Fusion module that aligns and integrates features in the frequency domain, and (2) an Audio-aware Aggregation module that performs advanced feature integration while considering the consistency between modalities. These strategies are embedded within a progressive fusion framework to enhance overall feature learning. Extensive experiments on the STARSS2023 dataset validate our dual-strategy approach, demonstrating that DSF-Net achieves state-of-the-art performance and outperforms existing methods. The source codes are publicly available at https://github.com/Devin-Pi/avseld-mamba.
Detecting audio-visual DeepFake (AVDeepFake) is becoming increasingly important as synthetic media tools become widely accessible and spread across consumer devices. In this study, we present an Detecting audio-visual DeepFakes (AV-DeepFakes) has become increasingly critical with the rapid proliferation of accessible synthetic media generation tools across consumer platforms. In this work, we propose a highperformance, deployment-efficient AV-DeepFake detection framework tailored for real-world consumer devices. The proposed model integrates a 3D convolutional visual encoder with a 2D convolutional audio encoder to learn synchronized multimodal representations, effectively capturing spatial, spectral, and prosodic inconsistencies inherent in manipulated content. To detect temporal forgeries, we introduce a bidirectional complementary boundary module that precisely localizes manipulation onsets and offsets. A cross-modal attention fusion mechanism aggregates modality-specific cues, while an uncertainty-aware gating strategy suppresses unreliable signals to improve robustness. Furthermore, a cross-modal discrepancy minimization loss encourages alignment for genuine samples while maximizing divergence for forged content, strengthening multimodal consistency learning. Extensive evaluations on FaceForensics++ and LAV-DF demonstrate the effectiveness of the proposed approach, achieving 97.1% AUC for clip-level detection and an 81.2% F1-score for temporal boundary localization, while reducing inference time by 5× compared with transformer-based methods.
Nasir Saleem, Adeel Hussain, Sami Bourouis et al.· International Journal of Int...· 0 citations
Audio-visual event recognition (AVER) has achieved significant performance improvements through transformer-based multimodal architectures. However, the high computational complexity, large memory footprint, and inference cost of these models hinder their deployment on edge and resource-constrained devices. This paper presents an efficient compression framework for hybrid cross-attention-based audiovisual event recognition by combining architectural model compression, knowledge distillation, and dynamic INT8 quantization. A high-capacity teacher model integrates VideoMAE for visual representation learning, the Audio Spectrogram Transformer (AST) for audio feature extraction, and a hybrid cross-attention fusion network for multimodal feature integration. A lightweight student model is constructed by reducing the hidden feature dimension, the number of attention heads, and the feedforward network size while preserving the overall network architecture. The student model is trained using knowledge distillation to effectively transfer discriminative knowledge from the teacher. Finally, dynamic INT8 post-training quantization is applied to further reduce the model size for efficient deployment. Experimental results on the Audio-Visual Event (AVE) dataset show that the proposed framework reduces the number of trainable parameters in the multimodal fusion module by 59.06%, with only a 2.14% decrease in classification accuracy compared with the teacher model. Furthermore, dynamic INT8 quantization reduces the model size from 10.71 MB to 2.04 MB while maintaining competitive recognition performance. These results demonstrate that the proposed framework provides an effective trade-off between recognition accuracy and computational efficiency, making it a promising solution for deployment on resource-constrained edge AI platforms.
Parinaz Binandeh Dehaghani, Danilo Pena, A. P. Aguiar· arXiv.org· 0 citations
The rapid advancement of voice synthesis technologies such as text-to-speech and voice conversion poses significant threats to speech-based authentication systems, necessitating robust deepfake detection methods. In this work, we propose a novel E-Branchformer-based architecture that effectively leverages self-supervised speech representations for audio deepfake detection. Our model employs parallel branches to simultaneously capture global contextual dependencies through multi-head self-attention and local temporal patterns through convolutional processing. To enhance discriminative capability, we integrate depthwise convolution and Squeeze-and-Excitation modules that enrich the classification token with refined patch token information after feature merging. Extensive experiments on ASVspoof 2021 LA, DF, and In-the-Wild datasets demonstrate state-of-the-art performance with equal error rates of 0.88%, 1.85%, and 6.30% respectively, substantially outperforming existing methods. Comprehensive ablation studies validate that the dual-branch architecture provides complementary discriminative information, Squeeze-and-Excitation Aggregation significantly improves SSL feature integration, and the combination of DWConv and SE modules is critical for effective class token enhancement. The superior performance on real-world scenarios demonstrates strong generalization capability to diverse acoustic conditions and unseen spoofing attacks.
Phuong Dat, Học Thủ, T. Nguyễn et al.· 0 citations
Event-based cameras provide a powerful sensing modality for capturing dynamic scenes with high temporal resolution and low redundancy. However, leveraging modern deep learning architectures often benefits from transforming asynchronous event streams into suitable intermediate representations. The Compact Spatio-Temporal Representation (CSTR) addresses this by encoding events into a three-channel image-like format compatible with standard vision backbones, but its reliance on mean timestamps limits its ability to model long and complex temporal dynamics. In this study, we propose a Generalized Compact Spatio-Temporal Representation (gCSTR), a simple yet effective extension of the CSTR that enhances temporal expressiveness by constructing complementary representations in the spatio-temporal $xt$ and $yt$ projection planes. These representations preserve the fine-grained temporal structure while maintaining compatibility with conventional convolutional neural networks. To effectively combine multiple gCSTR representations, we introduce the Parallel Specialization Network (PaSNet), a multi-branch architecture. Each branch is trained on a distinct gCSTR view, allowing specialization to the structure of each representation. We evaluate gCSTR and PaSNet across a wide range of event-based object and action recognition benchmarks and demonstrate substantial improvements over the standard CSTR, particularly for long and complex action sequences. Our approach achieves state-of-the-art performance on several challenging benchmarks, while matching or closely approaching prior work on others, and provides new insights into when spatio-temporal projections are beneficial for event-based vision tasks.
Per Nyblom, David Gustafsson, Tomas Wilkinson· IEEE Access· 0 citations
Sound event localization and detection (SELD) aims to identify active sound event classes and estimate their spatial locations from audio signals. Stereo SELD remains challenging because two-channel recordings provide limited directional and distance information for jointly estimating sound event detection (SED), direction of arrival (DOA), source distance, and source coordinates. This paper proposes a Gated Multi-Head Network (GMHN) for DCASE 2025 Task 3 stereo SELD. The proposed model is built on a ResNet-Conformer backbone and uses an 8-channel semantic–spatial feature representation as the encoder input. In addition, task-specific auxiliary cues are incorporated into different prediction branches: energy cues are used to improve source distance estimation (SDE), phase cues are used to refine DOA estimation, and BEATs features are used to enhance SELD. For source coordinate estimation (SCE), the model combines a raw coordinate branch and a geometry-based coordinate branch derived from DOA and distance predictions through a learnable gate. This design encourages geometric consistency among direction, distance, and coordinate estimates while preserving the flexibility of direct coordinate regression. Experimental results show that the proposed model improves SELD performance by achieving more accurate distance and localization estimation.
Min-Jin Kim, Seok-Pil Lee· Electronics· 0 citations
Identifying the impact location (“sweet spot”) on a tennis racket is crucial for performance evaluation in tennis training. However, existing approaches typically rely on expensive vision-based systems or specialized sensors, limiting their applicability in real-world scenarios. We propose a sound-sensor-based multi-task framework for racket impact localization using acoustic signals, combining radial region classification with continuous position regression. To effectively model complex acoustic patterns, we design a multi-expert convolutional neural network (CNN) architecture with multi-scale feature extraction and task-specific optimization. Each expert branch operates at a different temporal receptive field and is trained with tailored loss functions, enabling complementary learning of global patterns, class imbalance characteristics, and hard samples. The shared backbone jointly supports both classification and regression tasks, allowing the model to learn more informative and structured representations. Experimental results demonstrate that the proposed framework consistently outperforms conventional methods in radial region classification while achieving accurate impact position estimation. Furthermore, additive noise augmentation significantly improves robustness, enabling stable performance under noisy and practical sensing conditions.
Shaochi Zhang, Xiaoai Wang, Xuan Chang et al.· Applied Sciences· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.