Skip to content
Open access

Comparative Evaluation of Machine Learning Models for Predictive Maintenance Using the AI4I 2020 Dataset: Effects of Data Balancing and Feature Selection

2026 · ITEGAM- Journal of Engineering and Technology for Industrial Applications (ITEGAM-JETIA) · 0 citations

TL;DR

The main conclusions obtained demonstrate that it is possible to build reliable models capable of maximizing fault detection and minimizing false alarms, which is vital for predictive maintenance systems in industries where accuracy is critical.

Abstract

This article presents a comparative analysis of machine learning (ML) models applied to machine fault classification, using the AI4I 2020 dataset. The performance of five algorithms (random forest, decision tree (J48), artificial neural networks (ANN), k-nearest neighbors (KNN), and Naive Bayes) is evaluated in six different data scenarios: synthetic minority oversampling (SMOTE) technique, undersampling, and original imbalanced distribution in the complete feature set and in the feature selection set. In addition, the impact of feature selection techniques on predictive power is evaluated. The goal is to identify the optimal predictive model and compare the performance of these approaches to determine the advantages of feature reduction and data balancing strategies. To ensure robust validation, the models are analyzed using precision, recall, specificity, accuracy, and F1 score. The main conclusions obtained demonstrate that it is possible to build reliable models capable of maximizing fault detection and minimizing false alarms, which is vital for predictive maintenance systems in industries where accuracy is critical.

Read PDF

Similar papers

Open access 2026

A Comparative Analysis of Machine Learning Algorithms for the Early Prediction of Diabetes with an Evaluation of Class-Imbalance Handling

The study comes to the conclusion that headline accuracy is an unreliable guide in imbalanced medical prediction, that imbalance handling can change a model's practical usefulness, and that this benefit is strongly algorithm-dependent, meaning that the decision to resample should be based on the algorithm and the scree...

A. Oduroye, Temilade Opanuga, Esther Tosin Akanbi et al. · 0 citations
Conference Open access 2026

Supervised Learning for Classification in Data Science: A Comparative Perspective

Overall, logistic regression demonstrates strong and good generalization ability, while ensemble methods constitute robust alternatives for binary classification on structured data.

Zahra Benider, H. Bouzahir, Jaafar Idrais · 0 citations
Open access Aug 2026

Enhancing Predictive Accuracy and Operational Performance Based on Entropy–TOPSIS Framework for Machine Learning Model Selection for Predictive Maintenance

In the context of Industry 4.0, predictive maintenance increasingly relies on machine learning models to anticipate equipment failures and reduce unplanned downtime. This makes the selection of the most suitable ML model a multidimensional and complex decision problem, since models with comparable predictive performanc...

Zouhair Marmoucha, Mohamed el Khaili, A. Soulhi et al. · 0 citations
Open access Sep 2026

Intelligent Predictive Analytics using Machine Learning: A Comparative Evaluation of Classification Algorithms for High-Dimensional Data

This study comparatively evaluated the predictive performance of selected machine-learning classification algorithms for high-dimensional data using a quantitative computational design-and-evaluation methodology. The analysis involved data preprocessing, feature processing, model development, hyperparameter optimizatio...

Muhammad Awais, Muhammad Haad, Hameed Hussain et al. · 0 citations
Open access Sep 2026

Performance Analysis of Ensemble Technique for Classifying Imbalanced Dataset Using SMOTE-TOMEK Links Sampling

This century has seen a substantial increase in the value of data. Many decisions are made based on data. More the data, more possibility to get accurate result. But in many cases, available datasets are imbalanced i.e. one class have more data and other class have few data which leads to inaccurate classification whil...

Ajaya Shrestha, Dhiraj Pyakurel, Amit Rauniyar · 0 citations
Open access Sep 2026

Machine Learning-Based Predictive Analytics: Investigating the Influence of Feature Selection and Model Optimization on Prediction Accuracy

This study comparatively evaluated the predictive performance of machine learning classification models and examined whether feature selection and model optimization improved classification accuracy and overall predictive effectiveness. A quantitative computational research design was employed using a structured datase...

Muhammad Ahsan Tariq, Fasie Haider, Taaha Shahzad et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.