Machine-learning-based intrusion detection systems can learn nonlinear and interaction-based traffic patterns that are difficult to capture using static rules, but their predictions remain difficult to interpret reliably in analyst-facing cybersecurity workflows. This paper proposes a unified quantitative framework for auditing the reliability of Explainable Artificial Intelligence (XAI) in intrusion detection systems. The framework combines SHAP attributions with permutation-based functional importance, SHAP-permutation rank agreement, sufficiency and comprehensiveness retraining tests, retraining-stability analysis, spurious-feature injection, and a composite XAI reliability radar profile. The framework is evaluated on UNSW-NB15 and IoT-ToN using XGBoost, Logistic Regression, and Random Forest. Final performance is reported on untouched test dataset after training-only cross-validation and model selection. On UNSW-NB15, Random Forest and XGBoost achieved comparable held-out performance, with ROC AUC, PR AUC, and F1-scores of 0.9862, 0.9934, and 0.9231 for Random Forest, and 0.9858, 0.9932, and 0.9217 for XGBoost. Logistic Regression performed lower, with scores of 0.9695, 0.9774, and 0.9037. On IoT-ToN, XGBoost and Random Forest achieved near-ceiling performance, with ROC AUC values of 0.9994 and F1-scores of 0.9867 and 0.9879, respectively, while Logistic Regression degraded substantially, with ROC AUC of 0.8502 and F1-score of 0.6081. The reliability results show that high predictive performance does not automatically imply trustworthy explanations. Logistic Regression produced the most stable explanations, but its weaker detection performance limited its practical suitability. Among the high-performing models, XGBoost provided the strongest balance on IoT-ToN, while Random Forest provided the strongest balance on UNSW-NB15. These findings demonstrate the need to evaluate XAI-enabled intrusion detection using both predictive metrics and quantitative explanation-reliability diagnostics.
Elijah M. Maseno, Yanxia Sun, Zenghui Wang· Journal of Computer Virology...· 0 citations
Deep learning based network intrusion detection systems (IDS) can achieve strong traffic classification performance, but their resilience to adversarial manipulation remains a critical concern. This study evaluates the adversarial robustness of Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) models in a multiclass intrusion detection setting using the Train_Test_Network dataset with ten traffic classes. The models were trained on true sliding flow-window sequences under a unified preprocessing pipeline to support fair comparison. Adversarial robustness was first assessed under a white-box Fast Gradient Sign Method (FGSM) setting and then broadened through additional FGSM and Projected Gradient Descent (PGD) stress testing. SHapley Additive exPlanations (SHAP) were further used to analyse explanation instability under clean and adversarial conditions, and explanation-drift features were evaluated as a secondary adversarial detection signal. Under clean evaluation, both models achieved strong and nearly identical performance, with accuracies of 0.9614 for LSTM and 0.9615 for GRU and weighted F1-scores of 0.9597 and 0.9598, respectively. Under the main FGSM condition, performance declined substantially: the LSTM achieved adversarial accuracy of 0.6094 and weighted F1-score of 0.6290 with an evasion rate of 37.38%, while the GRU achieved adversarial accuracy of 0.5130 and weighted F1-score of 0.5690 with an evasion rate of 47.02%. The broader robustness sweep showed that iterative PGD exposed stronger fragility than FGSM alone. SHAP analysis indicated that adversarial perturbation altered both prediction outcomes and local explanation structure. A learned explanation-driven detector improved over the rule-based baseline, while larger-scale validation confirmed that explanation drift remained informative, though not perfectly separable, at broader scale. Overall, the results show that strong clean performance does not imply adversarial robustness, and that explanation drift provides a useful auxiliary signal for adversarial monitoring in recurrent IDS models.
Elijah M. Maseno, Yanxia Sun, Zenghui Wang· International Journal of Inf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.