LFA-MFNet: Local Feature-Aligned Multisource Information Fusion for Fault Diagnosis Under Limited and Imbalanced Conditions
Abstract
Real-world data-driven fault diagnosis is constrained by scarce fault samples, limited and imbalanced (L&I) class distributions, and heterogeneous multisensor data under operating condition shifts, which impede stable and generalizable discriminative representation learning. We propose the local feature alignment multimodal fusion network (LFA-MFNet), an end-to-end framework that fuses vibration, acoustic, and electrical current signals. A dual-stream convolutional encoder extracts modality-specific features, while local alignment and adaptive fusion enhance fine-grained cross-modal correspondence. A cross-modal Transformer then models global dependencies to yield more discriminative fused representations. To stabilize training under L&I conditions, we use multilevel supervision, alignment and consistency constraints, and domain-adversarial learning to improve cross-condition generalization. At the output stage, a parallel coarse-to-fine prediction strategy mitigates label-bias-induced errors. Experiments on two incremental imbalanced benchmarks show that the LFA-MFNet outperforms representative baselines in overall accuracy, minority-class recognition, and cross-condition generalization, and ablations verify local alignment and adaptive fusion as key contributors.