An integrated evaluation protocol for adversarial robustness, generalization, and explanation stability in URL-based phishing detection
The reliability of phishing Uniform Resource Locator (URL) detectors under adversarial URL rewriting, domain shift, and explanation instability remains insufficiently understood. This study proposes an integrated robustness evaluation protocol for URL-based phishing detection, which integrates structured adversarial perturbation, unseen attack-family generalization, compositional attack effects, explanation stability, and external vulnerability transfer. The protocol tests four representative model families: Logistic Regression, XGBoost, CharCNN, and BERT-base, using 235,370 validated URLs from PHIUSIIL, consisting of 100,520 phishing and 134,850 benign URLs, along with 49,121 PhishTank-validated phishing URLs for external validation. All models performed well on the clean test sets, ranging from 0.9962 to 0.9984, but their robustness decreased substantially under realistic URL mutations. Subdomain injection degraded the strong performance of Logistic Regression, XGBoost, and CharCNN to around 0.432, indicating collapse to the phishing-prevalence floor. BERT was highly susceptible to homoglyph, padding, and path-based perturbations. Leave-one-family-out evaluation also showed poor transfer to unseen subdomain attacks for both Logistic Regression and XGBoost, with Robustness Degradation Index values of 0.535 and 0.565, respectively. Explanation stability also suffered, with SHAP top-K Jaccard similarity dropping to 0.526-0.535 under subdomain perturbation. These results provide a solid benchmark for evaluating robustness-aware phishing URL detection for achieving deployable reliability under realistic adversarial and non-IID settings.