Skip to content
Open access

VA-DFN: An acoustic-vibration collaborative fusion network for bearings in strong noise environments

Jul 2026 · Measurement and control (London. 1968) · Vol 59, pp. 1324 - 1342 · 0 citations · 33 references

TL;DR

The VA-DFN demonstrates exceptional noise-resistant robustness under varying signal-to-noise ratio (SNR) conditions from −6 dB to 2 dB, achieving a maximum diagnostic accuracy of 99.55%, which is significantly superior to existing single-modality and conventional deep learning baseline models.

Abstract

To address the limitation that a single sensor is insufficient for comprehensively extracting deep fault features in strong industrial noise environments, which constrains bearing diagnosis accuracy, this paper proposes an acoustic-vibration collaborative fusion network. First, an Adaptive Gated Residual Block (AGRB) is designed and combined with a Twin-Gated Residual Block (TGRB) architecture to effectively extract highly robust deep local acoustic and vibration features amidst strong background noise. Second, a Bidirectional Attention Sensing Module (BASM) is constructed to perform deep interaction and complementary calibration of heterogeneous acoustic-vibration features in the global semantic dimension, breaking through the limitations of traditional shallow concatenation of multimodal features. To verify the effectiveness of the proposed model, an experimental study was conducted on a 6205 deep groove ball bearing using a non-contact acoustic-vibration synchronous acquisition system with a 25 cm acoustic monitoring distance and a 5096 Hz sampling rate. The dataset contains nine diagnostic categories, including one healthy state and eight fault states.Experimental results indicate that this method can achieve deep dynamic alignment of heterogeneous data. The VA-DFN demonstrates exceptional noise-resistant robustness under varying signal-to-noise ratio (SNR) conditions from −6 dB to 2 dB, achieving a maximum diagnostic accuracy of 99.55%, which is significantly superior to existing single-modality and conventional deep learning baseline models.

Read PDF

Similar papers

Open access Aug 2026

Acoustic–vibration fusion bearing fault diagnosis via a multi-scale Swin–CNN hybrid architecture

An acoustic–vibration fusion method for bearing fault diagnosis based on a multi-scale Swin–CNN hybrid architecture that employs a Bayesian optimization-based tunable Q-factor wavelet transform (BO-TQWT) to enhance fault-sensitive subbands under low signal-to-noise ratio conditions, and converts acoustic and vibration signals into two-dimensional time–frequency maps.

Mengran Liu, Zhao-Tao Du, Zhen-Xiang Xiong et al. · 0 citations
Aug 2026

PMMDA based on the fusion of acoustic and vibration signals under time-varying speed conditions for bearing fault diagnosis

In engineering applications, mechanical equipment must adapt to complex and dynamic working environments, where the rotational speed often varies over time, resulting in significant distribution discrepancies across different operating conditions. Meanwhile, information obtained from a single vibration signal is often insufficient and susceptible to external interference. Traditional single-source domain adaptation methods may suffer from negative transfer and fail to effectively exploit complementary knowledge from multiple source domains for target-domain fault diagnosis, resulting in reduced reliability and generalization performance of diagnostic models. To address these limitations, this paper proposes a Progressive Multi-Dimensional Multi-Source Domain Adaptation (PMMDA) method. From the perspective of collaborative utilization of multi-source data, the proposed method integrates multimodal information from vibration and acoustic signals and employs a multi-level feature alignment strategy to achieve progressive alignment between source and target domains. Additionally, an adaptive weighting mechanism is introduced to dynamically balance the contributions of different source domains during model training, thereby enhancing the overall learning performance. Experimental results on two sets of bearing fault diagnosis tasks under time-varying rotational speed conditions demonstrate that the proposed method can effectively mitigate the impact of distribution discrepancies, significantly improving the accuracy and generalization capability of the diagnostic model, and verifying its potential and reliability in complex operating conditions.

He Qin, Zhongwei Zhang, Xinyu Li et al. · 0 citations
Open access Aug 2026

Dual-Domain Fusion Network for Multi-Event Recognition in Φ-OTDR Sensing Systems

Leveraging advances in artificial intelligence algorithms, distributed acoustic sensing (DAS) based on phase-sensitive optical time-domain reflectometry (Φ-OTDR) has achieved high event-recognition accuracy through a variety of learning models. Nevertheless, further improving the accuracy of multi-event recognition remains a persistent challenge. In this paper, we propose a Dual-Domain Fusion Network (DD-FusNet) for vibration event recognition in Φ-OTDR sensing systems. To fully capture signal dynamics, the model simultaneously processes time- and frequency-domain representations, employing a crucial cross-attention mechanism to bridge these branches and enable dynamic, learnable interactions. Experimental results based on a six-class field engineering vibration event dataset collected by Φ-OTDR, containing car events, manual tapping, road breaker, excavation, leaking and noise, demonstrate that the proposed method achieves an average accuracy of 99.12%, significantly outperforming baseline methods by approximately 3 to 10 percentage points in accuracy, thereby ensuring the accuracy of multi-event recognition. We believe the proposed DD-FusNet will advance the recognition capabilities of Φ-OTDR systems in complex industrial sensing applications.

Rong Wang, Xinlei Qian, Chongyi Huang et al. · 0 citations
Open access Sep 2026

DS-GCMAF: A dual-stream gated cross-modal attention fusion network for robust bearing fault type and severity classification under complex noise

Accurate assessment of fault severity in rolling bearings under strong noise remains a critical challenge for intelligent predictive maintenance. In this study, the diagnostic task is formulated as joint fault type and severity classification, with emphasis on severity-level discrimination. In practical industrial environments, vibration signals are frequently contaminated by complex noise, including Gaussian noise, pink noise, Laplace (impulsive) noise, and combined interference, which severely mask transient fault impulses and distort time–frequency representations. To tackle this problem, we propose a Dual-Stream Gated Cross-Modal Attention Fusion Network (DS-GCMAF). The framework simultaneously processes one-dimensional (1D) raw vibration sequences and two-dimensional (2D) time–frequency representations. A bidirectional cross-modal multi-head attention mechanism is introduced to facilitate effective information interaction and alignment across heterogeneous feature spaces. Meanwhile, a quality-aware adaptive gating strategy is employed to dynamically regulate the contribution of each modality based on its reliability. In low signal-to-noise ratio (SNR) conditions, the less reliable 2D branch is selectively attenuated, and the more robust 1D stream serves as an anchor for cross-modal calibration. Extensive experiments were carried out on the Paderborn University (PU) dataset and HUSTbearing dataset under four complex noise types. On the PU dataset across varying SNR levels (-8 dB to 8 dB), DS-GCMAF achieves 85.80% accuracy at -8 dB, surpassing the strongest baseline by 2.69 percentage points. On the HUSTbearing dataset under -10 dB, the proposed method attains 74.04% accuracy under the most challenging combined noise condition, significantly outperforming existing single-stream and conventional fusion approaches. These results demonstrate the superior robustness and its effectiveness in joint fault-type and severity classification under highly noisy environments.

Yu He, Yuxuan Liu, Nai-Quan Su et al. · 0 citations
Open access Jul 2026

Multimodal vibro-acoustic bearing-fault diagnosis method based on channel fusion and the VACF-CNN model

Bearing-fault diagnosis based on single-modal signals is often constrained by incomplete fault information, whereas existing multimodal fusion methods often suffer from feature redundancy and a heavy preprocessing burden. To overcome these limitations, a vibro-acoustic channel fusion convolutional neural network model that employs channel-level fusion within an end-to-end one-dimensional framework, enhanced by wavelet packet decomposition and a squeeze-and-excitation mechanism, is proposed in this study. This method achieves mean diagnostic accuracies of 99.77% and 98.98% on a self-collected dataset and on the public MAFAULDA dataset, respectively. Additional experimental results indicate that the model maintains strong diagnostic performance under additive white Gaussian noise and variable-load conditions.

Keqin Ding, An Sun, Min Cao et al. · 0 citations
Open access Jul 2026

A multi-level constrained vibration–acoustic multimodal contrastive learning for cross-machine motor fault diagnosis

A novel vibration–acoustic multimodal contrastive learning framework designed to jointly regularize vibration–acoustic features at the levels of sample distribution, feature statistics, and feature structure enhances multi modal consistency, feature discriminability, and information diversity.

Yuan Zhuang, Deqiang He, Zhen-Zhen Jin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.