Dynamic Weighting and Adaptive Sparse Transformer for Federated Fault Diagnosis
Abstract
With the rapid development of the Internet of Things (IoT) and edge computing, the scale and complexity of modern networks have increased significantly, driving the demand for distributed fault diagnosis. Federated learning (FL) effectively addresses the issues of data privacy and dispersion by enabling edge devices to collaboratively train models without sharing raw data. However, existing FL-based fault diagnosis methods still encounter the following challenges. Firstly, static aggregation strategies struggle to balance the contributions of heterogeneous clients dynamically. Secondly, traditional local models are unable to effectively decouple sparse high and low-frequency features in fault signals, thereby limiting the accuracy of fault identification. Finally, the resource constraints of edge devices restrict the deployment of complex diagnostic models. To address these challenges, we propose a federated learning hybrid dynamic weight adjustment method based on delay and model quality, introducing the concept of “accelerated depreciation” in accounting and taxation and the concept of “asset allocation” in economics to improve the communication efficiency of fault diagnosis and reduce the impact of delay differences due to device heterogeneity on the effect of fault diagnosis. Additionally, we propose an adaptive sparse low high frequencies Transformer, introducing a lightweight attention mechanism and an adaptive feature extraction layer, which significantly reduces the computational overhead while maintaining high diagnostic accuracy. The experimental results show that, compared with the most competitive baseline, our method improves the fault diagnosis accuracy by 1% on the Case Western Reserve University Bearing dataset (CWRU), 0.78% on the Xi’an Jiaotong University Gearbox dataset (XJTU), and 0.44% on the Micro service Edge Computing dataset (MICRO).