Skip to content

Author

Zenghui Wang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Reliability auditing of explanations for machine-learning-based intrusion detection systems

Machine-learning-based intrusion detection systems can learn nonlinear and interaction-based traffic patterns that are difficult to capture using static rules, but their predictions remain difficult to interpret reliably in analyst-facing cybersecurity workflows. This paper proposes a unified quantitative framework for auditing the reliability of Explainable Artificial Intelligence (XAI) in intrusion detection systems. The framework combines SHAP attributions with permutation-based functional importance, SHAP-permutation rank agreement, sufficiency and comprehensiveness retraining tests, retraining-stability analysis, spurious-feature injection, and a composite XAI reliability radar profile. The framework is evaluated on UNSW-NB15 and IoT-ToN using XGBoost, Logistic Regression, and Random Forest. Final performance is reported on untouched test dataset after training-only cross-validation and model selection. On UNSW-NB15, Random Forest and XGBoost achieved comparable held-out performance, with ROC AUC, PR AUC, and F1-scores of 0.9862, 0.9934, and 0.9231 for Random Forest, and 0.9858, 0.9932, and 0.9217 for XGBoost. Logistic Regression performed lower, with scores of 0.9695, 0.9774, and 0.9037. On IoT-ToN, XGBoost and Random Forest achieved near-ceiling performance, with ROC AUC values of 0.9994 and F1-scores of 0.9867 and 0.9879, respectively, while Logistic Regression degraded substantially, with ROC AUC of 0.8502 and F1-score of 0.6081. The reliability results show that high predictive performance does not automatically imply trustworthy explanations. Logistic Regression produced the most stable explanations, but its weaker detection performance limited its practical suitability. Among the high-performing models, XGBoost provided the strongest balance on IoT-ToN, while Random Forest provided the strongest balance on UNSW-NB15. These findings demonstrate the need to evaluate XAI-enabled intrusion detection using both predictive metrics and quantitative explanation-reliability diagnostics.

Elijah M. Maseno, Yanxia Sun, Zenghui Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.