Skip to content

Author

Ariel Misael Orellana Albarracin

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Sep 2026

Comparative Analysis of Transformer-Based and ClassicalMachine Learning Models for Phishing Email Detection:A Multi-Source Dataset Evaluation with Explainability

Phishing remains one of the most persistent cyber threats, particularly in email environments where deceptive messages can be distributed at scale. This paper compares five classifiers: Multinomial Naive Bayes, Random Forest, Bidirectional Long Short-Term Memory (BiLSTM), DistilBERT, and BERT-base. A multi-source corpus of 82,689 cleaned and deduplicated emails was built from nine public datasets. Under a unified protocol, BERT-base achieved the highest F1-score (0.9824), while DistilBERT obtained an almost identical F1-score (0.9822) with lower measured inference latency (1.130 versus 2.218 ms/email), representing the strongest accuracy–latency trade-off in the evaluated environment. LIME explanations exposed plausible phishing indicators, such as urgency and account-verification language, but also mixed local contributions that require cautious interpretation. In the source-held-out experiment, the positive-class prevalence changed from 41.1% in training to 26.6% in testing, and DistilBERT produced 505 false negatives but only two false positives. Consequently, recall decreased from 0.9750 to 0.5346, showing that high mixed-source test performance does not guarantee robustness when complete data sources are unseen.

Andre Sebastian Samaniego Buñay, Ariel Misael Orellana Albarracin, Joel Marcelo Chuquimarca Pomagualli · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.