Predicting Quotation Success Based on Historical Project Data in Aluminum SMEs Using Classification Algorithms
Abstract
Small and medium enterprises (SMEs) in Indonesia's B2B aluminum and glass construction sector convert only 13–15% of their quotations, which means most of the pre-bid work on site assessment, costing, and proposal preparation produces no revenue. Machine learning has performed well at predicting B2B sales outcomes in large enterprises, but its usefulness for SMEs working with small, imbalanced datasets and few available features has received little empirical attention. This study benchmarks nine machine learning configurations on 300 historical quotation records from an Indonesian aluminum SME, where 14.3 % of quotations were won. The configurations cover Logistic Regression, Support Vector Machine, Random Forest, and XGBoost, each tested with and without enhanced imbalance handling, plus a soft voting ensemble. The methodology uses leak-safe SMOTE and BorderlineSMOTE oversampling, cost-sensitive learning, and nested 10-fold stratified cross-validation with grid-search hyperparameter tuning. The best configuration was Logistic Regression with BorderlineSMOTE and class-weight balancing, which achieved an F1-score of 0.309 ± 0.087 and an AUC-ROC of 0.661 ± 0.145, marginally above the more complex tree-based and ensemble alternatives. All nine configurations fell within a narrow AUC band of 0.579 to 0.661, which suggests that the predictive ceiling is determined by the features available rather than by the choice of algorithm. The practical implication is that SMEs in this position should invest in data infrastructure before algorithm selection. The study also offers a reproducible methodology for similar resource-constrained B2B settings.