Machine learning-based fault detection in low-speed bearings using a multi-environmental dataset
Abstract
This study provides a systematic robustness evaluation of classical machine learning for vibration-based bearing fault detection in low-RPM internal combustion engines (1000–2000 RPM) across a controlled temperature × humidity grid, a regime underrepresented in benchmark datasets that emphasise high-speed applications. A publicly available dataset from a 658cc engine (–10 °C–45 °C, 0%–100% humidity) was analysed; vibration features were derived from the non-zero channels of a tri-axial acquisition, with the bearing-housing vibration carried primarily by channel Ch3. To prevent temporal leakage, 390 263 continuous measurements were aggregated into 89 steady-state units, each spanning 90 s, yielding a deliberately independence-preserving but low sample-to-feature ratio (89:92). Four algorithms Random Forest, Support Vector Machine, Logistic Regression, and Neural Network were evaluated using stratified 5-fold cross-validation. All models achieved apparent accuracy exceeding 95%, with Random Forest performing best (97.8% ± 2.7%), but no statistically significant differences were found (Friedman test, p = 0.732). Vibration features, particularly crest factor and root mean square, provided the greatest discriminative power, while environmental factors accounted for less than 17% combined importance. The near-perfect linear separability (98.2% with Logistic Regression) indicates that the dataset’s binary, controlled-laboratory labelling rather than intrinsic bearing-degradation physics drives the clean classification. Accordingly, the reported accuracies are apparent upper-bound estimates from an exploratory study, not expected field performance; validation on 500–1000 or more samples with progressive-degradation labelling is essential before any operational claim can be supported.