Skip to content

Machine learning-driven fault detection in rotating machinery: A predictive maintenance framework

2026 · Materials Research Proceedings · 0 citations

TL;DR

A new predictive maintenance model that incorporates multi-domain feature extraction, and hybrid feature selection approach and optimized ensemble classifier to facilitate high-fidelity fault detection in four operational modes is proposed, affirming its suitability for deployment in industrial cyber-physical monitoring systems.

Abstract

Abstract. The traditional condition monitoring methods, such as threshold-driven vibration warning systems and periodic check-up programs, do not have the discriminative ability needed to differentiate between incipient faults signatures and normal operational variation. This paper proposes a new predictive maintenance model that incorporates multi-domain feature extraction, and hybrid feature selection approach and optimized ensemble classifier to facilitate high-fidelity fault detection in four operational modes, which are normal operation, inner-race bearing fault, outer-race bearing fault, and gear tooth spalling. Vibration signals are acquired at 12 kHz using tri-axial accelerometers, pre-processed through adaptive band-pass filtering and Z-score normalization, and subsequently transformed into a 47-dimensional feature space spanning time-domain statistics (RMS, kurtosis, skewness, crest factor), fast Fourier transform (FFT) spectral amplitudes, and discrete wavelet transform (DWT) energy coefficients. A two-stage feature selection pipeline—combining mutual information ranking with recursive feature elimination reduces the feature space to 18 optimal descriptors. An ensemble of random forest with Bayesian-optimization in cross validation of 10 folds and stratified cross-validation reach a mean accuracy of 98.7 percent, which is statistically better than support vector machine (SVM) and Artificial Neural Network (ANN) baselines by 4.1 and 2.9 percentage points respectively. The proposed framework demonstrates strong generalisability across variable operating speeds and load conditions, affirming its suitability for deployment in industrial cyber-physical monitoring systems.

View source

Similar papers

Review Open access Jul 2026

Enhanced Rotating Machinery Fault Diagnosis Using Holo-Hilbert Spectrum Analysis and Machine Learning

Reliable fault diagnosis in rotating machinery is challenging due to the nonlinear and non-stationary nature of vibration signals. Although time–frequency analysis is widely used, it cannot capture the cross-scale coupling between amplitude-modulated (AM) and frequency-modulated (FM) components that carry essential diagnostic information. This study applies Holo-Hilbert Spectrum Analysis (HHSA) to extract amplitude–frequency modulation features and integrates them with six machine learning classifiers to identify four fault conditions. Random Forest, K-Nearest Neighbors, and Logistic Regression achieve accuracies of up to 99.95%, yielding higher accuracy than Fast Fourier Transform-based features. The proposed framework employs an HHSA-based feature extraction pipeline that effectively captures AM–FM coupling in nonlinear vibration signals. It also provides higher discriminative capability than traditional spectral approaches and maintains robustness across multiple classifiers. This method offers high diagnostic accuracy and strong potential for industrial predictive maintenance. Future work will focus on improving computational efficiency and evaluating the framework under more diverse and realistic operating conditions. Received: 10 September 2025 | Revised: 20 April 2026 | Accepted: 10 June 2026 Conflicts of Interest The authors declare that they have no conflicts of interest to this work. Data Availability Statement The VBL-VA001 datasets that support the findings of this study are openly available at https://doi.org/10.1007/s42417-023-00959-9, reference number [44]. Author Contribution Statement Van-Trung Nguyen: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Data curation, Writing - original draft, Writing - review & editing, Visualization. Ba-Tan Le: Investigation, Data curation, Writing - original draft, Writing - review & editing. Van-Phuong Dao: Methodology, Validation, Writing - review & editing.

Van-Trung Nguyen, Ba-Tan Le, van-Phuong Dao · 0 citations
Conference Jul 2026

Machine Learning Algorithms for Bearing Fault Diagnosis Using Time-Domain Features

Rolling Element Bearing (REB) failures represent a major challenge in the maintenance of rotating machinery, as they compromise the reliability and operational continuity of systems. The analysis of vibration signals at the bearing level provides an effective approach for the early detection of anomalies and their classification, thereby helping to anticipate breakdowns and enhance equipment safety. In this work, the the well-known bearing dataset of Case Western Reserve University (CWRU) is utilized to perform fault classification using four machine learning algorithms: K-Nearest Neighbors (KNN), Decision Tree (DT), Random Forest (RF), and Gradient Boosting Decision Tree (GBDT). The study particularly addresses feature extraction in the time domain, such as root mean square (RMS), standard deviation, kurtosis, and other statistical indicators.

Ines Jaffel, Yasmine Atala, H. Messaoud · 0 citations
Conference Jul 2026

Multi-Sensor Fusion and Frequency-Domain Analysis for Predictive Maintenance of Industrial Induction Motors

Unplanned failures of induction motors impose serious operational and financial penalties on industrial facilities, yet the fault signatures that precede such failures are detectable well in advance through careful sensor instrumentation and data-driven analysis. This paper presents an end-to-end Internet-of-Things (IoT) predictive maintenance scheme based on two off-the-shelf sensors: a DS18B20 one-wire digital thermometer and 2 piezo vibration sensor modules, with an ESP32 edge device for running the full machine learning pipeline offline, independent from any cloud services. Four operating scenarios are considered: healthy condition, BPFO (bearing outer race fault), misaligned shaft, and rotor imbalance. Based on 1-second sampling intervals, 17 descriptors are derived, including statistics in the time domain, Fourier harmonic peaks, energy ratios between different frequency bands, and temperature gradients measured across all sensors. A dual-stage feature selection method using mutual information (MI) score and Random Forest mean decrease impurity (MDI) ranking reduces the number of features to the 10 most relevant descriptors, reducing the computational complexity by 41% at the expense of 5.6% F1-macro. On a balanced 600-sample synthetic dataset, the resulting Random Forest classifier attains 87.3% hold-out accuracy, 91.0±1.9% five-fold cross-validation accuracy, and a macro area-under-the-ROC-curve of 0.980. End-to-end inference takes just 39 ms on the ESP32, easily meeting the 200 ms requirement for real-time alerting.

Akash Mastud, Dhiraj Vaidya, Azaroddin Sayyed et al. · 0 citations
Open access Aug 2026

Comparative Evaluation of Machine Learning Algorithms for Fault Diagnosis in Automotive Press Lines

Minimizing unplanned downtime is critical for maintaining productivity in modern manufacturing. While combining sensor networks with machine learning provides a practical way to detect mechanical failures early, conventional data-driven diagnostics often fail during highly transient stamping operations. This failure stems from severe spectral smearing and signal distortions caused by fluctuating process loads and variable operating speeds. To address these limitations, we present a field-tested fault diagnosis (FD) framework deployed in an active automotive components plant. Over a twelve-month observation period, we collected raw vibration and process data from two operational transfer presses, building a comparative dataset that captures both localized gear damage and healthy baseline dynamics. After preprocessing the data to isolate signal anomalies, we systematically evaluated the diagnostic performance of six algorithms: SVM, Random Forest, Naive Bayes, k-NN, Decision Trees, and Logistic Regression. By integrating angle-based position data from a high-resolution encoder, the developed framework successfully pinpointed specific defective gear teeth. Ultimately, the Random Forest model outperformed the others, delivering the most robust detection accuracy under real-world factory conditions. These results show that the proposed Condition Monitoring (CM) approach significantly reduces resource waste and prevents costly downtime, offering a practical and scalable asset management model for industrial applications.

Ahmet Erdem Oner, Meral Bayraktar · 0 citations
Open access Jul 2026

Mode entropy knowledge machine: a fully automated bearing fault diagnosis model for complex operating conditions

Accurate diagnosis of rolling bearing faults is critical to the reliability of industrial equipment. However, rolling bearings often operate under complex operating conditions, and with data imbalances and noise interference, fault diagnosis of them remains extremely challenging. To address these issues, a novel mode entropy knowledge machine (MEKM) framework for robust bearing fault diagnosis is proposed in this study. For MEKM, the mode entropy space is firstly constructed to decompose the vibration signal into intrinsic mode components, and the noise-resistant feature extraction and dimensionality reduction are realized by principal component analysis. Secondly, a fast classifier based on extreme learning machines is introduced, and its parameters are automatically adjusted through a particle swarm optimization to establish an adaptive extreme learning machine diagnosis model, ensuring optimal generalization under different load and speed levels. Then, a collaborative optimization paradigm is developed to coordinate mode entropy features and classifier parameters through fully automated learning, in which entropy-driven feature characterization guides the iterative refinement of decision boundaries, while classifier feedback dynamically improves the selectivity of entropy features. Finally, validation is performed on bearings with multiple operating conditions, and the results indicated that the MEKM outperformed conventional deep learning methods in terms of diagnostic accuracy and generalization ability. The work provides a theoretical basis and an industrially feasible solution for health monitoring of mechanical equipment.

Hongchuang Tan, Yiheng Su, Jiang Ding et al. · 0 citations
Jul 2026

Machine learning-based fault detection in low-speed bearings using a multi-environmental dataset

This study provides a systematic robustness evaluation of classical machine learning for vibration-based bearing fault detection in low-RPM internal combustion engines (1000–2000 RPM) across a controlled temperature × humidity grid, a regime underrepresented in benchmark datasets that emphasise high-speed applications. A publicly available dataset from a 658cc engine (–10 °C–45 °C, 0%–100% humidity) was analysed; vibration features were derived from the non-zero channels of a tri-axial acquisition, with the bearing-housing vibration carried primarily by channel Ch3. To prevent temporal leakage, 390 263 continuous measurements were aggregated into 89 steady-state units, each spanning 90 s, yielding a deliberately independence-preserving but low sample-to-feature ratio (89:92). Four algorithms Random Forest, Support Vector Machine, Logistic Regression, and Neural Network were evaluated using stratified 5-fold cross-validation. All models achieved apparent accuracy exceeding 95%, with Random Forest performing best (97.8% ± 2.7%), but no statistically significant differences were found (Friedman test, p = 0.732). Vibration features, particularly crest factor and root mean square, provided the greatest discriminative power, while environmental factors accounted for less than 17% combined importance. The near-perfect linear separability (98.2% with Logistic Regression) indicates that the dataset’s binary, controlled-laboratory labelling rather than intrinsic bearing-degradation physics drives the clean classification. Accordingly, the reported accuracies are apparent upper-bound estimates from an exploratory study, not expected field performance; validation on 500–1000 or more samples with progressive-degradation labelling is essential before any operational claim can be supported.

P. Pugazhendi, Vinoth Vishwanathan, Aadil Arshad Ferhath et al. · 0 citations