Skip to content
Open access

Prior-enhanced cross-modal vibration and acoustic fusion network based on directed attention for machine fault diagnosis

Aug 2026 · Measurement science and technology · Vol 37, pp. 356113 · 0 citations · 38 references
Physics

TL;DR

The results of two machine fault experiments demonstrate that the proposed prior-enhanced cross-modal vibration and acoustic data fusion network based on directed attention mechanisms can achieve a significant leading advantage in the same diagnostic tasks.

Abstract

Existing intelligent fault diagnosis methods based on multimodal fusion face the problem of significant differences in the representation capabilities of different modal data for machine faults, making it difficult to achieve optimal cross-modal data fusion and accurate fault identification. This study proposes a prior-enhanced cross-modal vibration and acoustic data fusion network based on directed attention mechanisms to address the aforementioned issue. First, through preliminary experiments in fault diagnosis, the differences in fault sensitivity between vibration and acoustic data are measured to determine the prior dominant data modality. Then, based on the directed cross-attention mechanism, a prior-dominant modality-weighted fusion of vibration and acoustic data features is realized. This process allows for unidirectional feature information transfer from the dominant data modality to the weaker one, avoiding reverse information contamination. Thus, cross-modal fusion features that are more sensitive to machine faults can be extracted. Finally, the extracted cross-modal fusion features are used to achieve fault diagnosis. The results of two machine fault experiments demonstrate that, compared with the state-of-the-art methods, the proposed method can achieve a significant leading advantage in the same diagnostic tasks.

Read PDF

Similar papers

Open access Aug 2026

Research on Fault Location Technology of Transformer Acoustic Feature Recognition and Field Perception Data Fusion under Small Sample Parameters

Reliable transformer fault diagnosis under limited fault samples remains a significant challenge in intelligent power systems. To address the difficulties associated with weak fault signatures, severe environmental interference, and insufficient training samples, this study investigates transformer fault location technology based on acoustic feature recognition and field perception data fusion. The generation mechanism and propagation characteristics of transformer acoustic signals are first analyzed, and an improved time–frequency feature extraction method is developed to enhance feature representation under small-sample conditions. A multi-physics data fusion framework integrating acoustic, vibration, and electrical sensing information is then established, and a dedicated attention mechanism is designed to achieve deep feature fusion across heterogeneous data sources. Finally, an enhanced deep neural network model is employed for accurate fault localization and condition identification. Experimental results demonstrate that the proposed framework effectively improves fault recognition performance and location accuracy under small-sample constraints. The study provides technical support for intelligent power equipment monitoring and offers methodological references for signal propagation analysis, sensor fusion, and electromagnetic condition monitoring systems.

Y.-G. Li, L.-J. Feng, R.-R. Li et al. · 0 citations
Open access Aug 2026

Multimodal Heterogeneous CNN with Adaptive Modality Fusion for Intelligent Fault Diagnosis of Bearings

In industrial equipment fault diagnosis, vibration and acoustic signals are highly complementary yet exhibit significant differences in frequency distribution and noise sensitivity. Traditional multimodal methods generally rely on homogeneous feature extractors and direct feature concatenation, which may fail to capture modality-specific characteristics and introduce irrelevant information during fusion. To address this, we propose a novel multimodal heterogeneous convolutional neural network framework. Specifically, separate 1D CNN branches are designed for vibration and acoustic signals. Their architectural differences are determined by the characteristics of each sensing modality. The vibration branch focuses on extracting high-level discriminative fault features, including impulse responses and modulated components from vibration signals, while the acoustic branch is designed to preserve fragile high-frequency details of acoustic signals. Furthermore, an adaptive cross-attention fusion module is introduced to dynamically model cross-modal dependencies, assigning Softmax-based weights to enhance dominant features and suppress noise. Experiments based on bearing fault experimental data demonstrate that the proposed heterogeneous architecture significantly outperforms traditional homogeneous models. The dynamic weighting mechanism effectively prevents inferior noisy modalities from degrading overall performance, achieving high diagnostic accuracy. Although validated on rolling bearing fault diagnosis, the proposed heterogeneous multimodal framework is not restricted to bearings and can be readily extended to other intelligent condition monitoring tasks involving heterogeneous sensor fusion, such as gearboxes, motors, and other rotating machinery.

Chang Sun, Chenkun Wang, Shiwei Huang et al. · 0 citations
Open access Aug 2026

A gear fault diagnosis method based on neural network architecture search with cross-modal dynamic convolution enhancement

Multimodal fault diagnosis has attracted widespread attention in the field of rotating machinery due to the complementary information provided by multiple sensor signals. However, the design of signal fusion networks and the extraction of correlation features between multiple signals remain challenging. To address this, this paper proposes a multimodal framework based on differentiable architecture search to fuse vibration and acoustic emission signals, two signals of vastly different orders of magnitude, and to achieve automatic search for the fusion network. Simultaneously, a cross-modal dynamic convolutional module is introduced to extract and fuse correlation features between multiple signals. Experimental results on two benchmark gear fault datasets demonstrate that, compared to traditional models and existing network search methods, the proposed method exhibits stronger robustness under different operating conditions and achieves superior diagnostic performance. Ablation experiments further validate the effectiveness of the proposed interaction mechanism and joint optimization strategy.

Jialin Li, Kun Long, Yong Xiao et al. · 0 citations
Open access Jul 2026

Prior-augmented measurement-signal fusion using a dual-branch cross-attention network for cross-domain bearing fault diagnosis

Cross-machine bearing fault diagnosis is strongly affected by inconsistencies in vibration measurement conditions, including rotational speed, sampling frequency, and structural transmission paths. These factors cause speed-induced fault-frequency drift and cross-domain distribution discrepancies, making features learned from source-domain measurements unreliable in target-domain scenarios. Existing transfer learning (TL) methods are predominantly data-driven and insufficiently exploit mechanism-related prior information, which limits their interpretability and cross-machine generalization. To address these challenges, this paper proposes a prior-augmented dual-branch cross-attention network, termed PA-DCA Net, for cross-domain adaptive bearing fault diagnosis. First, a multi-scale S-transform with channel-weighted fusion is used to construct informative time-frequency representations from vibration signals. Meanwhile, a speed-normalized prior feature is introduced at the input-feature level to encode the relative rotational-speed discrepancy between the current operating condition and the source-domain reference condition. This prior feature is combined with eight conventional time-domain statistical features to form a statistical-prior feature vector. Second, an image-statistical dual-branch network is constructed, in which the image branch extracts deep time-frequency features and the statistical branch maps the statistical-prior vector into a high-dimensional representation. Multi-head cross-attention is then employed to achieve directed feature interaction between the two modalities. Third, a progressive TL framework integrating source-domain supervised pretraining, few-shot target-domain fine-tuning, CORAL, multi-kernel maximum mean discrepancy, and FixMatch-based consistency regularization is adopted. The proposed method is validated on five cross-domain tasks constructed from three public bearing datasets. PA-DCA Net achieves average accuracies of 97.10% and 96.92% on the Case Western Reserve University (CWRU)-to-Jiangnan University and CWRU-to-Huazhong University of Science and Technology cross-machine tasks, respectively, outperforming several representative transfer-learning baselines.

Xinyu Zhu, Hua Huang, Xilong Zhang et al. · 0 citations
Open access Aug 2026

Multimodal gated fusion and domain adaptation for cross-condition high-speed train bearing fault diagnosis

A four-branch multi-modal unsupervised domain-adaptive fault diagnosis framework, termed CRG-DA Net, based on ConvNeXt and ResNet1D, aimed at enabling cross-condition fault diagnosis under unlabeled target data is proposed.

Zhihao Zhao, Li Xu, Jingjing Cai et al. · 1 citation
Jul 2026

Cross-Attention Fusion of Time-Frequency Acoustic Features for Insulation Fault Recognition in Power Equipment Using a Microphone Array System.

This protocol provides a robust non-contact strategy for insulation fault diagnosis and condition monitoring of electrical power equipment by extracting and fusing time- and frequency-domain acoustic features for automated fault classification.

Weifeng Chen, Chunguang Hou, Yu Gu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.