Explainable AI for Intrusion Detection on Live Multi-Service Honeypot Telemetry
Abstract
# Are Explanations Faithful Under Fire? Evaluating XAI Reliability on Live Honeypot Intrusion-Detection Telemetry ## Supplementary Data and Code — Version 3.0 [](https://doi.org/10.5281/zenodo.22727709) > **Authors:** Sridhar G, J B Simha, Rashmi Agarwal> **Affiliation:** School of Computer Science and Application / RACE, REVA University, Bengaluru, India> **Contact:** sridhargovardhan@race.reva.edu.in> **Version:** 3.0 (September 2026) --- ## What Changed from Version 2.0 to 3.0 | Change | v2.0 | v3.0 ||---|---|---|| **Primary faithfulness** | Retraining-based AUDAC | **Frozen-model** masking-based AUDAC (no retraining) || **AUDAC values** | RF 0.679, XGB 0.662, LGB 0.688 | **RF 0.352, XGB 0.348, LGB 0.346, MLP 0.442, LSTM 0.463** || **Flip rate** | Malicious only | **Both malicious AND benign** (trees: 100%/0%) || **Per-instance SHAP-LIME ρ** | Negative (−0.18, buggy) | **Positive** (RF 0.72, XGB 0.43, LGB 0.52, MLP 0.58) || **LOSO** | RF only | **All 4 sklearn models** || **Feature ablation** | dst_port_id only | **4 features** (dst_port_id, src_port_class, hour_of_day, ip_reputation_score) || **Multi-seed** | RF only (3 seeds) | **All 6 models** × 3 seeds || **Service inventory** | 10 services listed | **2,309 unique ports** documented || **Transformer** | Negative control (F1 = 55.9%) | **Convergence failure** (F1 = 0.00 on 2/3 seeds) || **Labels** | "True labels" | **"Operational annotations"** || **Scripts** | 6 scripts | **10 scripts** || **Result files** | 2 CSVs | **7 CSVs + 2 TXT** | --- ## Overview This repository contains the complete reproducibility package for the manuscript *"Are Explanations Faithful Under Fire? Evaluating XAI Reliability on Live Honeypot Intrusion-Detection Telemetry."* The central contribution is a **frozen-model faithfulness evaluation** of model-agnostic explanations (SHAP, LIME) on live adversarial honeypot traffic. Each model is trained once, saved as a frozen checkpoint, and evaluated without retraining — directly measuring whether the trained model relies on the features SHAP identified. --- ## Repository Structure ```.├── README.md│├── # ── Primary Analysis: Frozen-Model Faithfulness ──├── frozen_model_faithfulness.py # Trains, saves, evaluates frozen models (6 models × 3 seeds)├── frozen_fixes.py # Corrects LIME name-matching bug + MLP SHAP Pipeline error├── frozen_faithfulness.csv # Primary results: 18 rows (6 models × 3 seeds)├── frozen_fixes_results.txt # Corrected SHAP-LIME ρ and MLP AUDAC│├── # ── Service Generalisation (Phase 4) ──├── phase4_generalisation.py # LOSO all models, multi-feature ablation, CIs├── loso_all_models.csv # LOSO F1 for RF, XGBoost, LightGBM, MLP (48 rows)├── multi_feature_ablation.csv # 4-feature ablation + combined removal (20 rows)├── per_service_all_models.csv # Per-service F1 for all 4 models (48 rows)├── service_inventory.csv # Complete inventory: 2,309 unique destination ports├── phase4_results.txt # Full Phase 4 output with bootstrap CIs│├── # ── Secondary Analysis: Retraining-Based ──├── run_faithfulness_retraining.py # Retraining-based AUDAC (secondary robustness analysis)├── xai_faithfulness_metrics.py # Faithfulness metrics library (AUDAC, flip rate, sparsity)├── faithfulness_retraining_results.csv # Retraining-based results (6 models, seed=42)│├── # ── Supporting Scripts ──├── per_service_evaluation.py # Per-service F1, LOSO, dst_port_id ablation (RF only, v2)├── bootstrap_ci.py # Bootstrap 95% CIs for AUDAC and flip rate├── reviewer_experiments.py # 6 reviewer-response analyses in one pass├── benchmark_diagnostics.py # Leakage audit and de-duplication analysis├── unsw_corrected_split_rerun.py # Corrected UNSW-NB15 official split re-run│├── # ── Benchmark Results ──└── verified_benchmark_results.csv # UNSW-NB15 and CIC-IDS2017 (6 models)``` --- ## Key Results ### 1. Primary Analysis: Frozen-Model Faithfulness (Table 9) Models trained once → saved as frozen checkpoints (.pkl/.keras) → top-k SHAP features masked with training-set medians → **no retraining**. This directly measures whether the trained model depends on the features SHAP identified. | Model | F1 | Frozen AUDAC ↓ | Mono. | Flip Mal | Flip Ben | Global ρ | Per-inst ρ ||---|---|---|---|---|---|---|---|| RF | 88.6% | **0.352** | Strict | 100% | 0% | 0.87 | 0.72 || XGBoost | 88.2% | **0.348** | Essential | 100% | 0% | 0.93 | 0.43 || LightGBM | 88.1% | **0.346** | Essential | 100% | 0% | 0.85 | 0.52 || MLP | 87.7% | **0.442** | Mostly | 98.8% | 2.2% | 0.87 | 0.58 || LSTM | 83.4% | **0.463** | Mostly | 87.0% | 13.6% | N/A | N/A || Transformer | 67.0% | 0.670 | N/A | 0% | N/A | N/A | N/A | **Monotonicity tiers:** Strict = no violations (k=0–10); Essential = violations only near chance; Mostly = one real violation, overall downward; Anti = accuracy rises (Transformer on retraining — excluded from primary). **Bootstrap 95% CIs (frozen, seed=42):**- AUDAC: RF [0.338, 0.366], XGBoost [0.332, 0.365], LightGBM [0.330, 0.362], MLP [0.425, 0.460]- Flip-rate: RF [1.000, 1.000], XGBoost [1.000, 1.000], LightGBM [1.000, 1.000], MLP [0.970, 1.000] **Key finding:** Tree ensembles flip 100% of malicious predictions but 0% of benign predictions when top SHAP features are masked — confirming that the identified features are specifically diagnostic of malicious behaviour. ### 2. Multi-Seed Variance (seeds 42, 7, 2024) | Model | F1 (mean ± SD) | AUDAC (mean ± SD) ||---|---|---|| RF | 0.891 ± 0.005 | 0.341 ± 0.018 || XGBoost | 0.884 ± 0.003 | 0.428 ± 0.086 || LightGBM | 0.888 ± 0.006 | 0.378 ± 0.027 || MLP | 0.881 ± 0.004 | 0.677 ± 0.024 || LSTM | 0.851 ± 0.014 | 0.461 ± 0.018 || Transformer | 0.223 ± 0.387 | 0.223 ± 0.387 | Transformer achieves F1 = 0.00 on two of three seeds — catastrophic convergence instability on tabular data. ### 3. SHAP-LIME Cross-Method Consistency (corrected in v3) | Model | Global ρ | Per-instance median ρ | IQR ||---|---|---|---|| RF | 0.87 | 0.72 | [0.63, 0.82] || XGBoost | 0.93 | 0.43 | [0.30, 0.59] || LightGBM | 0.85 | 0.52 | [0.38, 0.68] || MLP | 0.87 | 0.58 | [0.48, 0.68] | **Bug fix in v3:** Per-instance values in v2 were negative (median ρ = −0.18) due to a LIME feature-name matching error. LIME with `discretize_continuous=True` returns names like `"payload_length_bytes <= 245.50"` which did not match the original feature names during rank-correlation computation. Corrected in `frozen_fixes.py` using substring matching. ### 4. Service Generalisation (Phase 4 — new in v3) **Service inventory:** 2,309 unique dst_port_id values observed during the 21-day deployment:- 10 deliberately deployed honeypot services (ports 21, 22, 25, 53, 80, 443, 3306, 3389, 5900, 6379) — 651–744 sessions each- 2 additional ports with substantial traffic (8080 HTTP-alt, 8443 HTTPS-alt)- ~2,297 ephemeral scanning/probing ports (1–2 sessions each) **LOSO — all 4 sklearn models (12 services with ≥10 test samples):** | Model | Mean F1 Drop | Services Improved ||---|---|---|| RF | −0.10 pp | 4/12 || XGBoost | +0.54 pp | 4/12 || LightGBM | −0.07 pp | 6/12 || MLP | −0.43 pp | 7/12 | **Multi-feature ablation (all models):** | Feature Removed | RF | XGB | LGB | MLP ||---|---|---|---|---|| dst_port_id | −0.35 pp | −0.51 pp | −0.38 pp | −1.42 pp || src_port_class | +0.61 pp | +0.89 pp | +0.14 pp | +0.63 pp || hour_of_day | +0.81 pp | +1.51 pp | +0.38 pp | +0.96 pp || ip_reputation_score | −0.74 pp | +0.80 pp | +0.25 pp | −0.46 pp || **ALL FOUR** | **+2.80 pp** | **+3.36 pp** | **+2.68 pp** | **+5.43 pp** | Negative drop = F1 *improves* when that feature is removed. Removing dst_port_id improves F1 for all models — no shortcut learning. Removing all four service-related proxies costs only 2.7–5.4 pp, confirming that the core detection signal resides in behavioural features. ### 5. Secondary Analysis: Retraining-Based Faithfulness Retained from v2 as a robustness check. For each removal step k, the model is retrained on remaining features — allowing adaptation, producing higher (more conservative) AUDAC. | Model | F1 | AUDAC ↓ | Flip % | SHAP-LIME ρ ||---|---|---|---|---|| RF | 88.59% | 0.634 | 100% | 0.93 || XGBoost | 88.27% | 0.617 | 98.2% | 0.89 || LightGBM | 88.17% | 0.633 | 100% | 0.86 || MLP | 87.01% | 0.676 | 99.6% | 0.87 || LSTM | 83.15% | 0.630 | 12.8% | 0.93 || Transformer | 59.16% | 0.603 | 15.0% | 0.70 | ### 6. Benchmark Comparison (Table 10) | Model | UNSW-NB15 F1 | UNSW-NB15 AUC | CIC-IDS2017 F1 | CIC-IDS2017 AUC ||---|---|---|---|---|| XGBoost | 92.31% | 98.50% | 99.82% | 99.98% || RF | 92.24% | 98.29% | 99.80% | 99.98% || LightGBM | 92.06% | 98.49% | 99.90% | 99.99% || LSTM | 91.32% | 98.03% | 98.45% | 99.96% || MLP | 91.04% | 97.96% | 98.84% | 99.94% || Transformer | 90.77% | 96.86% | 97.27% | 99.92% | Note: different feature schemas (42 features for UNSW-NB15, 78 for CIC-IDS2017 vs 13 for the honeypot), different label definitions, and different preprocessing pipelines prevent direct comparison. Lower live-traffic scores may reflect adversarial difficulty but should not be interpreted as a causal relationship. --- ## Script Descriptions ### `frozen_model_faithfulness.py` — Primary Faithfulness (NEW in v3) The central experiment addressing the peer-review requirement for frozen-model evaluation. For each of 6 models × 3 seeds:1. Trains the model on the 85% training partition (n = 9,265)2. Saves a frozen checkpoint to `frozen_models/` (.pkl for sklearn, .keras for TensorFlow)3. Loads the frozen checkpoint for all subsequent evaluations4. Computes SHAP values (TreeExplainer for trees, KernelExplainer for deep models)5. Runs masking-based AUDAC (k = 0, 1, ..., 13) — **no retraining**6. Computes flip rate for **both** malicious and benign