Intelligent Predictive Analytics using Machine Learning: A Comparative Evaluation of Classification Algorithms for High-Dimensional Data
Abstract
This study comparatively evaluated the predictive performance of selected machine-learning classification algorithms for high-dimensional data using a quantitative computational design-and-evaluation methodology. The analysis involved data preprocessing, feature processing, model development, hyperparameter optimization, and stratified 5-fold cross-validation using Logistic Regression, Support Vector Machine (SVM), Random Forest, and XGBoost. Model performance was assessed using accuracy, precision, recall, F1-score, ROC-AUC, confusion matrices, and computational efficiency, while the Friedman test was used to examine differences in model performance across cross-validation folds. The results indicated that XGBoost achieved the highest predictive performance, with an accuracy of 96.27%, precision of 96.3%, recall of 95.8%, F1-score of 96.0%, and ROC-AUC of 0.982, followed by Random Forest with an accuracy of 94.18%, SVM with 91.36%, and Logistic Regression with 87.42%. XGBoost also produced the lowest false-positive and false-negative counts (18 and 21, respectively), whereas Logistic Regression required the least training time (1.84 seconds). The Friedman test indicated statistically significant differences among the algorithms for accuracy (χ²=12.84, p=.005), precision (χ²=11.76, p=.008), recall (χ²=12.31, p=.006), F1-score (χ²=12.57, p=.006), and ROC-AUC (χ²=13.42, p=.004). Overall, the findings identify XGBoost as the most effective algorithm for reliable predictive analytics in high-dimensional data environments, while demonstrating that model selection should balance predictive performance with computational efficiency.