Skip to content
Review Open access

Artificial Intelligence for Phishing and Social-Engineering Detection: A Machine Learning Evaluation on a Benchmark URL Dataset

Sep 2026 · International Journal of Emerging Engineering and Technology · 0 citations · 24 references

Abstract

Phishing remains the leading initial-access vector in modern cyberattacks, and the emergence of large language models (LLMs) has made social-engineering content easier to produce and harder to distinguish from legitimate communication. This paper reviews recent (2023-2026) research on artificial-intelligence-based phishing detection and reports an original empirical evaluation on a public benchmark dataset of 88,647 labelled URLs described by 111 engineered features spanning URL, domain, directory/file, query-parameter, and network/WHOIS-derived attributes. Five supervised learning models logistic regression, linear support vector machine, random forest, gradient boosting, and a multilayer perceptron were trained and evaluated using stratified five-fold cross-validation. The random forest classifier achieved the strongest performance (accuracy = 0.970, F1 = 0.957, ROC-AUC = 0.995), consistent with the ensemble-dominance trend reported in the recent literature. A follow-up robustness simulation shows that although the model is highly resistant to random feature-level noise, this does not imply resistance to the targeted, semantically-aware evasion strategies enabled by generative AI, which recent studies show can bypass commercial filters. The paper concludes with a discussion of the dual-use nature of AI in this domain and directions for more adversarial-robust, content-aware detection systems.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.