Skip to content
Conference

A Context-Aware NLP Model for Automated Spam Message Identification

Aug 2026 · 2026 International Conference on Secure Information Systems and Technologies (ICSIST) · pp. 1330-1337 · 0 citations · 17 references

Abstract

Spam communications remain a persistent cybersecurity threat, consuming network bandwidth, reducing productivity, and delivering malicious content such as phishing links and malware. Traditional rule-based and keyword-centric filtering methods struggle against modern spam that employs contextual manipulation and obfuscated text. This paper presents a context-aware Natural Language Processing (NLP) model for automated spam identification evaluated across heterogeneous communication corpora, including the SMS Spam Collection, Enron Email, and SpamBase datasets. The proposed system preprocesses text through cleaning, tokenization, stop-word removal, stemming, and TF-IDF vectorization. Five supervised machine learning algorithms (Naive Bayes, Logistic Regression, K-Nearest Neighbors, Support Vector Machines, and Random Forest) alongside fine-tuned transformer architectures (BERT and RoBERTa) are evaluated. Model predictions are validated using Explainable AI (XAI) frameworks, specifically SHAP and LIME, to interpret feature contributions. Furthermore, adversarial robustness tests evaluate model resilience against character substitution, word obfuscation, and URL manipulation. Experimental results indicate that Random Forest achieves an accuracy of 97.1% on traditional features, while fine-tuned BERT achieves superior contextual generalization across message and email datasets with an F1-score of 98.4%.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.