Skip to content
Open access

A Unified and Interpretable Benchmark of Classification Models for DGA-Based Power Transformer Fault Diagnosis

Sep 2026 · Applied Sciences · 0 citations

Abstract

Power transformers are among the most critical components of electric power transmission and distribution systems, and unexpected failures can lead to substantial economic losses and prolonged power outages. Dissolved Gas Analysis (DGA) is the most widely used diagnostic technique for transformer fault diagnosis and involves interpreting gases dissolved in insulating oil. However, conventional interpretation methods, such as the Rogers ratio, the Doernenburg ratio, the IEC 60599 ratio method, and the Duval Triangle, rely heavily on expert knowledge, may produce inconsistent diagnoses for the same oil sample, and may fail to provide a diagnosis in certain cases. In this study, 12 classification models were evaluated using the publicly available Power Transformers Fault Detection and Diagnosis (FDD) and Remaining Useful Life (RUL) dataset and a unified evaluation protocol. Model performance was assessed using Accuracy, Balanced Accuracy, and Macro-F1 score, while model interpretability was investigated through Shapley Additive Explanations (SHAP) analysis. The results showed that ensemble tree-based methods achieved the best overall performance. LightGBM and Random Forest both attained an Accuracy of 0.969 and a Macro-F1 score of 0.924, while LightGBM further achieved the highest Balanced Accuracy of 0.940. XGBoost exhibited the most stable performance under cross-validation. SHAP analysis revealed that engineered relative concentration features, particularly the CO/H2 ratio and the combined gas ratio, were among the most influential features for fault classification. These findings demonstrate that, for datasets of this scale, ensemble tree-based models combined with well-designed features provide strong and interpretable performance for imbalanced DGA-based fault diagnosis, highlighting the effectiveness of feature-based ensemble learning for small- to medium-sized datasets.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.