A Multi-Dataset Evaluation of Classical ML for SMS Spam Detection: Balancing Accuracy, Efficiency, and Generalizability
Abstract
SMS remains among the top popular modes of communication, both personally and professionally. However, recent concerns about the security aspect of mobile communications due to SMS spam/smishing messages have emerged. While Deep Learning shows promising results in text classification applications, its massive computing and memory consumption renders it unfeasible for deployment on low-resources devices. This work leverages the existing databases namely “UCI SMS Spam Collection” and various modern smishing datasets to assess the potential of using lightweight classical ML techniques including Multinomial Naïve Bayes (MNB), SVM and LR for detecting SMS spams. Three feature extraction approaches namely BoW, TF-IDF and TF-IDF combined with KMeans Clustering were employed. Our experiments showed that while models perform fairly decently when assessed on the training set, their performance deteriorates significantly when tested on other sets. Further analysis suggested that MNB provides an optimal trade-off between speed, accuracy and model simplicity. Finally our results show that the performances we got from our models are statistically significant.