Comparative Analysis of SVM, XGBoost, LSTM, and BERT for Spam Email Classification
Abstract
The rapid growth of digital communication has been accompanied by a sustained increase in spam messages, which pose risks ranging from user annoyance to phishing, fraud, and malware distribution. Prior work has largely evaluated either classical machine learning or deep learning models in isolation, and the studies that do compare them typically differ in dataset, preprocessing, and evaluation protocol, which makes their results difficult to compare directly. This study addresses that gap by benchmarking two classical machine learning models (Support Vector Machine and XGBoost) and two deep learning models (LSTM and BERT) on the same public dataset of 5,572 labeled messages (4,825 ham, 747 spam), using an identical preprocessing pipeline, an identical 80/20 train-test split, and the same evaluation metrics. BERT achieved the highest accuracy at 99.28%, followed by SVM at 98.65%, LSTM at 98.57%, and XGBoost at 97.67%. The advantage of BERT is concentrated in spam recall (0.96 against 0.91 for SVM and LSTM), indicating that contextual representations mainly reduce missed spam rather than false alarms. Notably, a linear SVM over TF-IDF features matched the deep learning models to within 0.7 percentage points at a small fraction of their training cost, indicating that the accuracy advantage of transformer models on this corpus is real but modest and must be weighed against a substantially higher computational budget.