Machine Learning-Based Anomaly Detection in Hydrogen Fuel Cells: A Path Toward Sustainable and Reliable Energy Systems
Abstract
To encourage the use of hydrogen energy and improve the dependability of hydrogen fuel cells, this study develops a two-stage semi-supervised framework for anomaly pattern discovery and codification. Isolation Forest, One-Class Support Vector Machine, and Local Outlier Factor are first applied to unlabeled time series data; six supervised classifiers are subsequently trained on ensemble-derived pseudo-labels to learn these patterns efficiently. The analysis uses a reproducible 20% sample (random seed 42) of 185,721 observations from four fuel cells in the NASA Prognostics Data Repository, yielding 37,144 observations. The three unsupervised algorithms exhibit distinct detection patterns, with 280 common detections (15.07% of each model’s flagged observations and 0.754% of the sample). A consensus-weighted ensemble identifies 1296 potential anomalies (3.49%). Random Forest best reproduces the ensemble pseudo-labels, with 99.365% testing accuracy; this value measures pattern learnability and is not independent verification of physical faults. Gini importance identifies power difference, capacity, load power, load current, measured power, and load resistance as the principal predictive variables. Because independently confirmed fault labels and maintenance outcomes are unavailable, the reported detections are interpreted as systematic deviations requiring electrochemical or expert verification.