A Comparative Analysis of Machine Learning Algorithms for Urban Traffic Congestion Prediction: The Impact of Hyperparameter Optimization
Abstract
Accurately predicting urban traffic congestion is a central requirement in modern Intelligent Transportation Systems (ITS). Although machine learning (ML) remains the primary tool for this task, model performance is strongly influenced by the characteristics and quality of the training data. This study presents a comparative evaluation of eight supervised ML algorithms Logistic Regression (LR), Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Gaussian Naïve Bayes (GNB), Gradient Boosting (GB), and the Multi-Layer Perceptron (MLP) Neural Network using a balanced synthetic dataset representing multiple congestion scenarios. Model behavior is assessed under both default configurations and after hyperparameter optimization using a lightweight Grid Search CV procedure. Several classifiers, including DT, RF, and GB, achieved exceptional accuracy exceeding 0.9977 percent, indicating strongly separable class boundaries within the dataset. Hyperparameter tuning proved particularly beneficial for SVM and KNN, substantially improving their generalization capabilities, while producing minimal changes in already well-optimized models such as Naïve Bayes and GB. The tuned RF emerged as the most reliable and robust classifier, offering high predictive accuracy with reduced overfitting risk. The study also highlights a critical consideration: models trained exclusively on clean, noise-free data may exhibit inflated performance that does not fully reflect real-world operational conditions. These findings underscore the importance of evaluating model robustness when deploying ML-based congestion-prediction systems in dynamic and noisy traffic environments.