Skip to content
Open access

Spam Email Classification Using TF-IDF and Classical Machine Learning on the Enron-Spam Corpus

Sep 2026 · Wasit Journal of Computer and Mathematics Science · 0 citations · 21 references

Abstract

Unsolicited email is a persistent operational and security burden and many high-performing neural approaches require computational and deployment costs that are not necessary for resource-constrained filtering systems. This paper tries to fill this gap by providing a rigorous and reproducible comparison of lightweight classifiers to discriminate spam from legitimate email without overestimating the corpus as a dedicated phishing or malware benchmark. Using 30,089 cleaned and de-duplicated messages from the Enron-Spam corpus, word-level unigram and bigram term frequency–inverse document frequency (TF-IDF) features were evaluated with Multinomial Naive Bayes, Logistic Regression, Linear Support Vector Machine, and Random Forest classifiers. Hyperparameters were selected by grid search on the training partition; stability was assessed through repeated stratified five-fold cross-validation across five random seeds; and paired Logistic Regression and Linear SVM predictions were compared using an exact two-sided McNemar test. On the held-out test set, Logistic Regression achieved the highest accuracy (99.15%) and F1-score (99.13%), followed by Linear SVM (99.05% accuracy; 99.02% F1). Naive Bayes achieved 98.59% accuracy, whereas Random Forest reached 96.69% despite a spam recall of 99.55%. Repeated cross-validation produced identical rounded mean accuracy and macro-F1 for Logistic Regression and Linear SVM (99.08% ± 0.11%), and their held-out difference was not statistically significant (p = 0.2101). These findings support tuned linear TF-IDF models as accurate, stable, and computationally practical spam-filtering baselines, while cross-corpus validation remains necessary.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.