Explainable AI for phishing URL detection: a Bayesian-optimized stacking ensemble framework with SHAP-guided feature learning
Abstract
Phishing remains one of the most persistent and financially damaging threats facing modern organizations, with over 4.7 million incidents recorded in 2023 alone. Existing AI-based phishing detection frameworks are constrained by limited benchmarking scope, absent model interpretability, and insufficient statistical validation — three limitations that collectively restrict operational utility in real-world security environments. We present an explainable, end-to-end machine learning pipeline evaluated on a large public benchmark of 247,950 URLs described by 41 structural and lexical features. The pipeline integrates SHAP-driven feature selection (reducing 41 to 24 features via a 95% cumulative-signal rule), a systematic benchmark of 12 classifiers spanning seven algorithmic families, Bayesian hyperparameter optimization via Optuna TPE sampling (40 trials each for XGBoost and CatBoost), and a heterogeneous stacking ensemble combining Optuna-tuned XGBoost, CatBoost, Extra Trees, and Random Forest under a logistic-regression meta-learner. A four-layer statistical validation protocol — comprising a Friedman omnibus test, Wilcoxon signed-rank tests, paired t-tests, and Cohen's d effect sizes — was applied to five-fold cross-validation accuracy distributions to assess directional consistency, with the limited inferential resolution of five folds explicitly acknowledged. SHAP-driven selection reduced the feature space by 41.5% while retaining 95% of predictive signal. The stacking ensemble achieved 96.75% accuracy, 96.74% F1-score, and AUC of 0.9947, attaining the lowest Brier score among all 13 models (0.0246), indicating superior probability calibration. The Friedman omnibus test confirmed significant performance differences across models (χ 2 F = 59.84, p < 0.0001), and all 12 Wilcoxon pairwise comparisons yielded the minimum attainable p-value ( p = 0.0313), confirming the ensemble never lost a cross-validation fold against any baseline. Post-hoc SHAP analysis identified subdomain structure, URL length, and URL entropy as the dominant phishing indicators at both ensemble and base-learner levels. The co-leaders — the stacking ensemble and Extra Trees — demonstrate that rigorous, interpretable AI pipelines can advance phishing detection accuracy and transparency simultaneously. The framework's calibrated risk scores, threshold flexibility, and multi-level SHAP explainability support analyst-facing decision-making in security operations, while its leakage-free