Aug 2026· Measurement science and technology· Vol 37· 0 citations· 36 references
Physics
TL;DR
An acoustic–vibration fusion method for bearing fault diagnosis based on a multi-scale Swin–CNN hybrid architecture that employs a Bayesian optimization-based tunable Q-factor wavelet transform (BO-TQWT) to enhance fault-sensitive subbands under low signal-to-noise ratio conditions, and converts acoustic and vibration signals into two-dimensional time–frequency maps.
Abstract
To address the problems that weak impulsive features in acoustic–vibration signals are easily masked under strong noise, fault severities within the same fault category are difficult to distinguish, and existing fusion models show insufficient coordination between local details and global semantics, this paper proposes an acoustic–vibration fusion method for bearing fault diagnosis based on a multi-scale Swin–CNN hybrid architecture. The proposed method first employs a Bayesian optimization-based tunable Q-factor wavelet transform (BO-TQWT) to enhance fault-sensitive subbands under low signal-to-noise ratio conditions, and then converts acoustic and vibration signals into two-dimensional time–frequency maps. Subsequently, a Swin–CNN hybrid network is constructed, in which CNNSwinBridge performs a two-stage, single-pass cross-guided recalibration between CNN-derived local features and Swin-derived contextual features. Specifically, Swin features provide channel-wise semantic guidance for CNN features, whereas the recalibrated CNN features provide spatial texture guidance for Swin features. The module provides lightweight cross-architecture coordination without recurrent or iterative feedback. Experimental results on the BJTU-RAO and University of Ottawa bearing datasets show that the proposed method achieves average accuracies of 99.51% and 99.96%, respectively, under clean operating conditions. Further analysis indicates that BO-TQWT provides relatively limited performance gains under clean conditions, whereas it can more effectively enhance fault-sensitive features under strong noise, thereby improving the fine-grained discrimination of different severity levels within the same fault category.
The VA-DFN demonstrates exceptional noise-resistant robustness under varying signal-to-noise ratio (SNR) conditions from −6 dB to 2 dB, achieving a maximum diagnostic accuracy of 99.55%, which is significantly superior to existing single-modality and conventional deep learning baseline models.
Fan-Long Zhu, Jun-Yu Lai, Pei-Wen Lu et al.· Measurement and control (Lon...· 0 citations
In industrial equipment fault diagnosis, vibration and acoustic signals are highly complementary yet exhibit significant differences in frequency distribution and noise sensitivity. Traditional multimodal methods generally rely on homogeneous feature extractors and direct feature concatenation, which may fail to capture modality-specific characteristics and introduce irrelevant information during fusion. To address this, we propose a novel multimodal heterogeneous convolutional neural network framework. Specifically, separate 1D CNN branches are designed for vibration and acoustic signals. Their architectural differences are determined by the characteristics of each sensing modality. The vibration branch focuses on extracting high-level discriminative fault features, including impulse responses and modulated components from vibration signals, while the acoustic branch is designed to preserve fragile high-frequency details of acoustic signals. Furthermore, an adaptive cross-attention fusion module is introduced to dynamically model cross-modal dependencies, assigning Softmax-based weights to enhance dominant features and suppress noise. Experiments based on bearing fault experimental data demonstrate that the proposed heterogeneous architecture significantly outperforms traditional homogeneous models. The dynamic weighting mechanism effectively prevents inferior noisy modalities from degrading overall performance, achieving high diagnostic accuracy. Although validated on rolling bearing fault diagnosis, the proposed heterogeneous multimodal framework is not restricted to bearings and can be readily extended to other intelligent condition monitoring tasks involving heterogeneous sensor fusion, such as gearboxes, motors, and other rotating machinery.
A hybrid deep learning architecture is proposed for robust vibration-based fault diagnosis in industrial machinery by jointly modeling time-domain, frequency-domain, and temporal dynamics, enabling coherent cross-domain interaction and robust fault characterization.
Canan Taştimur· Information Technology and C...· 0 citations
A four-branch multi-modal unsupervised domain-adaptive fault diagnosis framework, termed CRG-DA Net, based on ConvNeXt and ResNet1D, aimed at enabling cross-condition fault diagnosis under unlabeled target data is proposed.
Zhihao Zhao, Li Xu, Jingjing Cai et al.· Measurement and control (Lon...· 1 citation
The results show that SAMACNN outperforms both classical and advanced methods on the two datasets, demonstrating strong robustness and generalization capability in complex variable-condition measurement environments.
DaXin Li, Hong Wang, Hai Xue et al.· Engineering Research Express· 0 citations
A robust and noise-resilient bearing fault diagnosis framework that integrates advanced signal processing with hybrid deep learning techniques is presented, demonstrating strong robustness and generalization capability.
Sujit Kumar, Manish Kumar, Bam Bahadur Sinha· International Journal of Dyn...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.