A Fidelity-, Stability-, and Agreement-Aware Evaluation of Explainable Intrusion Detection for the Internet of Medical Things (CICIoMT2024)
Abstract
Software and result data accompanying the study "A Fidelity-, Stability-, and Agreement-Aware Evaluation of Explainable Intrusion Detection for the Internet of Medical Things (CICIoMT2024)" by Hayder Mohsin Hammood Al-Darraji, Shahinda Mohamed Elkholy, and Haitham A. El-Ghareeb (Department of Information Systems, Faculty of Computers and Information, Mansoura University, Egypt). This archive contains the complete, reproducible analysis pipeline and all result files (tables and figures) underlying the manuscript. The study operationalizes the trustworthiness of post-hoc explanations algorithmically as three distinct dimensions — fidelity (a SHAP-guided feature-removal flip test against a random-removal control), stability (SHAP ranking agreement across retraining seeds), and cross-method agreement (SHAP, LIME, and Permutation Importance) — for four tree-based intrusion-detection models (Decision Tree, Random Forest, XGBoost, Gradient Boosting) trained on the CICIoMT2024 WiFi-and-MQTT benchmark (7,139,549 training and 1,612,117 test flows). All four models achieve near-perfect detection (attack recall ≈ 0.999; Benign recall 0.949–0.970; macro-F1 0.978–0.986) yet dissociate sharply on the three trust axes: explanations show a small but statistically significant, boundary-concentrated fidelity advantage over random removal (McNemar p < .0001), high retraining stability (Spearman 0.839–0.988, n = 20 seeds), but weak cross-method agreement (Spearman ≤ 0.372) that persists after ruling out LIME sampling stochasticity and a sample-size confound. The work proposes an Explanation Trust Scorecard as an exploratory pre-deployment screening aid — not a validated, externally calibrated trust metric. Contents: trust_eval_ciciomt2.py — the full analysis pipeline (Python 3.12; numpy, pandas, scipy, scikit-learn, xgboost, shap, lime, matplotlib), plus auxiliary and robustness scripts. results_paper2/tables/ — every result table as CSV, including per-instance backing data for the fidelity, size-matched-control, boundary-stratified, operational-metrics, and correlation-stability analyses. results_paper2/figures/ — all manuscript figures (ROC curves, trust scorecard, agreement scatter grid, correlation heatmap, feature-cluster diagram, and confusion matrices). requirements.txt — dependency version ranges. NEW_SCRIPTS_AND_OUTPUTS.txt — a script-to-output manifest. Every value is the direct output of the corresponding script run on the CICIoMT2024 official train/test split with the fixed random seeds documented in the manuscript's methods (Section 2.5). Per-seed retrained model caches are not included, as they are fully regenerable from the scripts and seeds. The CICIoMT2024 dataset itself (Dadkhah et al., 2024) must be obtained from its original source.