Toward Intelligent Fault Diagnosis in Renewable Energy Systems: Integrating Data-Driven and Model-Based Approaches Under Variable Environmental Conditions
Renewable energy systems, including solar photovoltaic arrays and wind turbines, operate under highly variable environmental and operating conditions. Factors such as changing irradiance, temperature fluctuations, wind variability, and component aging make Fault Detection and Diagnosis (FDD) particularly challenging. Therefore, developing reliable, accurate, and interpretable diagnostic methods is essential to ensure system efficiency, safety, and long-term operation. Traditional model-based approaches, which rely on physical system models, offer clear interpretability and solid theoretical foundations. However, their effectiveness can be limited by modeling inaccuracies and difficulties in capturing complex nonlinear behaviors. On the other hand, data-driven and Artificial Intelligence (AI) techniques have demonstrated strong capabilities in pattern recognition and fault classification, but often face challenges related to data dependence, limited transparency, and reduced robustness under unseen conditions. This paper provides a comprehensive and structured review of FDD techniques for renewable energy systems, covering model-based, signal-based, data-driven, and hybrid approaches. A unified perspective is presented to clarify the strengths, limitations, and application domains of each category. Particular attention is given to recent advances in hybrid methods that combine physical modeling and AI, including feature fusion, ensemble learning, attention-based models, and transfer learning. Moreover, advanced signal processing techniques are discussed for their role in extracting meaningful features from noisy and non-stationary data. Rather than ranking methods by headline accuracy, which has become saturated and is only weakly comparable across heterogeneous datasets, the review adopts a critical, deployment-oriented perspective that emphasizes cross-condition robustness, standardized benchmarking, and the constraints of real-world deployment. The review also highlights the growing importance of digital twin technology as a promising framework for next-generation FDD systems, enabling real-time monitoring, adaptive learning, and predictive maintenance. Furthermore, Explainable AI is explored as a key direction for improving the transparency and trustworthiness of AI-based diagnostic models. Finally, the paper identifies major challenges and open research issues, such as data scarcity, generalization among different operating conditions, computational efficiency, and system reliability. Future research directions are outlined toward developing more robust, adaptive, and interpretable FDD solutions that can operate effectively in dynamic and uncertain environments.
With the large-scale integration of new energy into power systems, the intermittency and volatility caused by the high penetration of renewable energy generation (such as wind power and photovoltaic power) require diagnostic systems to possess strong uncertainty-handling capabilities. Consequently, the complexity and uncertainty of power grid operation have increased significantly, posing unprecedented challenges to the safe, stable and economic operation of the power system. The traditional fault diagnosis and handling methods, which are based on fixed models and manual experience, can no longer adapt to the dynamic and complex operating characteristics of the new energy power grid, making it urgent to explore intelligent technical solutions. Traditional power grid fault diagnosis approaches rely mainly on expert experience and physical models. Model-based methods locate faults through state estimation and power flow calculation, whose accuracy heavily depends on model precision and parameter identification. However, under complex operating conditions such as high new energy penetration, grid topology changes, and frequent fluctuations in power supply and demand, establishing an accurate mathematical model that can cover all operating scenarios is extremely challenging—model mismatches often occur, leading to reduced fault diagnosis accuracy. Expert systems, on the other hand, integrate the operational experience of power grid engineers into rule bases, offering transparent reasoning processes that are easy to understand and verify. Yet, they suffer from inherent limitations: knowledge acquisition is time-consuming and labor-intensive, it is difficult to update rules in a timely manner with the iteration of grid technology, and they lack self-learning ability, making it impossible to adapt to new fault types and complex operating environments brought by new energy integration. When dealing with massive real-time data generated by the power grid (including new energy output data, load data, equipment monitoring data, and environmental data) and complex system environments, the following prominent problems frequently arise, which further restrict the efficiency and reliability of power grid operation and fault handling.
Rui-Ze Ji· Highlights in Science Engine...· 1 citation
The rapid expansion of renewable energy systems demands reliable fault detection and prediction to ensure operational efficiency and grid stability. This study presents a novel framework that integrates Extended Kalman Filter (EKF) state estimation with uncertainty-aware graph learning for photovoltaic (PV) array fault detection and localization. Raw sensor data are processed by the EKF to generate refined state estimates and uncertainty covariances for each PV module. These uncertainty measures dynamically modulate an attention-based graph construction module, enabling adaptive edge weighting that down-weights unreliable connections during noisy or transient conditions. The resulting dynamic graphs are analyzed by a temporal graph attention network to produce both node-level fault localization and global anomaly scores. The graph-construction, temporal-encoding, and prediction components were optimized jointly, while the EKF process and observation models and their noise covariances remained fixed after calibration. On the real-world dataset, it attains an AUC-ROC of 0.941 and F1-score of 0.918 for global detection, and a node-level F1-score of 0.865 with Exact Match Ratio of 0.738 for fault localization. The approach demonstrates strong robustness to sensor noise and transient faults by leveraging physical uncertainty to guide graph topology. This work offers a promising direction for reliable monitoring of large-scale PV systems and other sensor-rich energy infrastructures.
Saud Wasly, N. Abu-Hamdeh· Scientific Reports· 0 citations
Energy is a fundamental component of modern society, enabling technological progress and economic development. However, conventional energy generation has led to significant environmental challenges, making the transition toward renewable energy sources a necessity in the current century. In addition to sustainability, renewable energy systems must also be economically viable. The overall cost of such systems is typically divided into CAPEX, associated with investment, and OPEX, related to maintenance and operation. In this context, this thesis proposes data-driven methods to optimize both cost components in renewable energy systems.
First, a novel hierarchical-based clustering method is developed to improve CEM, which determine the CAPEX of future energy systems. The proposed algorithm determines the optimal number of clusters using a modified elbow method enhanced with a stopping criterion to avoid unnecessary iterations. This criterion works based on percentage variance and runtime to determine the number of clusters systematically. The method then employs a hybrid selection strategy combining Euclidean distance, medoid, and mean values to determine the most representative vector in each cluster. To evaluate the performance, it is compared with well-known clustering methods. The numerical results show that not only does the proposed approach select a more appropriate number of clusters with lower computational cost than other systematic techniques, but also outperforms benchmark methods in terms of accuracy, computational efficiency, and robustness.
Second, this thesis presents a window-based deep learning framework for failure diagnosis in PV plants, contributing to the reduction of OPEX. The method incorporates two types of windows: one capturing historical operational data and another containing previous diagnostic decisions. This design addresses the impact of weather variability, which can mask failures, particularly under cloudy conditions. Four deep learning models, CNN, LSTM, ConvLSTM, and GRU, are also evaluated across multiple failure scenarios, including isolated and combined faults. The results show that incorporating historical information improves diagnostic performance compared to conventional approaches, while increasing the window size does not necessarily improve accuracy. In addition, combining relevant failure types further improves detection performance.
Finally, this thesis addresses residential energy systems by developing a degradation-aware battery modeling approach for home energy management systems. Since battery operation leads to degradation and increased OPEX, accurate modeling of battery behavior is essential. A physics-based model implemented in PyBaMM is tuned using four swarm intelligence algorithms. The results indicate that the whale optimization algorithm achieves superior performance in terms of accuracy and computational cost. The tuned model is then integrated into the optimization framework, where Pareto-optimal solutions are evaluated based on battery state of health. The solution with the highest battery state of health is then selected as the optimal strategy.
RESUMEN
La energía es un componente fundamental de la sociedad moderna, ya que permite el progreso tecnológico y el desarrollo económico. Sin embargo, la generación convencional de energía ha provocado importantes desafíos medioambientales, lo que hace que la transición hacia fuentes de energía renovables sea una necesidad en el siglo actual. Además de la sostenibilidad, los sistemas de energía renovable deben ser también económicamente viables. El coste total de estos sistemas se divide típicamente en CAPEX, asociado a la inversión, y OPEX, relacionado con la operación y el mantenimiento. En este contexto, esta tesis propone métodos basados en datos para optimizar ambos componentes de coste en sistemas de energía renovable.
En primer lugar, se desarrolla un novedoso método de agrupamiento basado en estructuras jerárquicas para mejorar los CEM, que determinan el CAPEX de los sistemas energéticos futuros. El algoritmo propuesto determina el número óptimo de clústeres mediante un método del codo modificado, mejorado con un criterio de parada para evitar iteraciones innecesarias. Este criterio se basa en el porcentaje de varianza y el tiempo de ejecución para determinar sistemáticamente el número de clústeres. Posteriormente, el método emplea una estrategia híbrida de selección que combina la distancia euclídea, el medoide y los valores medios para determinar el vector más representativo en cada clúster. Para evaluar su rendimiento, se compara con métodos de agrupamiento ampliamente conocidos. Los resultados numéricos muestran que el enfoque propuesto no solo selecciona un número de clústeres más adecuado con menor coste computacional que otras técnicas sistemáticas, sino que también supera a los métodos de referencia en términos de precisión, eficiencia computacional y robustez.
En segundo lugar, esta tesis presenta un marco de aprendizaje profundo basado en ventanas para el diagnóstico de fallos en plantas fotovoltaicas, contribuyendo a la reducción del OPEX. El método incorpora dos tipos de ventanas: una que captura datos históricos de operación y otra que contiene decisiones diagnósticas previas. Este diseño aborda el impacto de la variabilidad meteorológica, que puede enmascarar fallos, especialmente en condiciones nubladas. Asimismo, se evalúan cuatro modelos de aprendizaje profundo (CNN, LSTM, ConvLSTM y GRU) en múltiples escenarios de fallo, incluyendo fallos aislados y combinados. Los resultados muestran que la incorporación de información histórica mejora el rendimiento del diagnóstico en comparación con los enfoques convencionales, mientras que el aumento del tamaño de la ventana no necesariamente mejora la precisión. Además, la combinación de tipos de fallo relevantes mejora aún más el rendimiento de detección.
Por último, esta tesis aborda los sistemas energéticos residenciales mediante el desarrollo de un enfoque de modelado de baterías que tiene en cuenta la degradación para sistemas de gestión energética doméstica. Dado que el funcionamiento de la batería conduce a su degradación y a un aumento del OPEX, un modelado preciso de su comportamiento es esencial. Un modelo basado en principios físicos implementado en PyBaMM se ajusta utilizando cuatro algoritmos de inteligencia de enjambre. Los resultados indican que el algoritmo de optimización de ballenas logra un rendimiento superior en términos de precisión y coste computacional. El modelo ajustado se integra posteriormente en el marco de optimización, donde se evalúan soluciones óptimas de Pareto en función del estado de salud de la batería. Finalmente, se selecciona como estrategia óptima la solución que presenta el mayor estado de salud de la batería.
The rapid expansion of renewable energy infrastructures has introduced significant challenges for monitoring system performance, ensuring regulatory compliance, and maintaining transparency in energy production and emissions reporting. Modern renewable energy systems generate large volumes of heterogeneous operational data through smart meters, sensor networks, supervisory control and data acquisition (SCADA) systems, and distributed generation platforms. Traditional monitoring approaches, which rely primarily on rule-based thresholds and periodic audits, often struggle to process such complex and dynamic data streams in real time. This paper proposes an artificial intelligence (AI)-driven intelligent monitoring framework designed to enhance operational oversight and sustainability monitoring in renewable energy systems. The proposed architecture integrates machine learning and anomaly detection techniques to analyze energy production data, detect abnormal operational patterns, and assess compliance with environmental and regulatory requirements. A design-science research methodology is adopted to develop and evaluate the framework using simulated renewable energy datasets representing solar and wind energy production scenarios. Experimental results demonstrate that the AI-based monitoring system significantly improves anomaly detection accuracy and reduces reporting delays compared with traditional rule-based monitoring methods. The proposed approach supports intelligent renewable energy infrastructure management by enabling proactive monitoring, improved operational transparency, and enhanced sustainability reporting.
Badreddine Said, Ashraf Rashid, Omari Asem et al.· E3S Web of Conferences· 1 citation
An experimental/synthetic hybrid, data-driven FDI framework that leverages supervised machine learning (ML) integrated with an experimentally validated second-order electro-thermal battery model to generate a mixed experimental–synthetic dataset enables fast, real-time diagnosis of complex multi-fault scenarios at the cell or module level in series–parallel LIB pack architectures.
Taha Mohamed Abdelatif Maaradji, Saïd Alem, Emanuele Gravante et al.· Transactions of the Institut...· 0 citations
Power transformers are critical components of modern power grids, and their operational reliability directly affects power system security, stability, and continuity. With the increasing intelligence and complexity of power systems, condition monitoring and fault diagnosis of transformers have received growing attention. However, conventional diagnostic methods often face limitations such as complex modeling procedures, high computational costs, weak adaptability, and insufficient generalization under nonlinear and coupled operating conditions. In recent years, surrogate models have emerged as effective tools for transformer fault diagnosis because of their advantages in high-dimensional nonlinear mapping, rapid prediction, and data-driven approximation. This paper systematically reviews the research progress of surrogate models in transformer fault diagnosis and establishes a classification framework from the perspectives of model types, modeling strategies, data sources, and application scenarios. The principles, applicable conditions, and performance characteristics of representative surrogate models are comparatively analyzed. Furthermore, the advantages and limitations of different surrogate modeling approaches are discussed in terms of diagnostic accuracy, stability, generalization ability, interpretability, and computational efficiency. Although surrogate models show strong potential for intelligent transformer fault diagnosis, challenges remain in small-sample learning, data quality dependence, model interpretability, and cross-condition generalization. Future research should focus on multi-source information fusion, integration of physical mechanisms with data-driven learning, lightweight intelligent modeling, and standardized evaluation systems. This review aims to provide methodological guidance and technical references for the development of reliable, efficient, and interpretable transformer fault diagnosis methods.
Guang-Fen Wan, Kai Yang, Fei Xiong et al.· Italian National Conference...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.