Modelling Healthy Operations of Generator Bearings in Wind Farms
This study investigates the consistency and impact of feature selection methods on deterministic and probabilistic normal behaviour models (NBMs) for estimating rear generator bearing temperature in wind farms using SCADA data. Rigorous NBM development enables the transfer of key features, enhances anomaly detection comparability and facilitates potential portability across wind farms. The rise of machine learning and a vast range of statistical methods available for anomaly detection or estimation with SCADA data further emphasizes the need for investigating consistent NBMs. This paper evaluates four feature selection methods—Pearson's correlation coefficient (PCC), decision tree weights (DTW), mutual information (MI) and Shapley values of a neural network (SHAP)—for two wind farms to assess consistency in identified influential features and their impact on estimation algorithms. The results highlight the importance of prioritizing features consistently identified across datasets and methods over merely optimizing deterministic estimation performance. Temperature‐related sensors dominated the key features, but their specific locations were also ranked consistently. Probabilistic interval estimations demonstrated superior estimation performance and to identify whether a significant difference in feature can be attributed to the model or to an anomaly in the measured data.