Skip to content
#explainable ai Dataset Open access

Explainable AI for Intrusion Detection on Live Multi-Service Honeypot Telemetry

Sep 2026 · Zenodo (CERN European Organization for Nuclear Research)

Abstract

# Are Explanations Faithful Under Fire? Evaluating XAI Reliability on Live Honeypot Intrusion-Detection Telemetry ## Supplementary Data and Code — Version 3.0 [![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.22727709.svg)](https://doi.org/10.5281/zenodo.22727709) > **Authors:** Sridhar G, J B Simha, Rashmi Agarwal> **Affiliation:** School of Computer Science and Application / RACE, REVA University, Bengaluru, India> **Contact:** sridhargovardhan@race.reva.edu.in> **Version:** 3.0 (September 2026) --- ## What Changed from Version 2.0 to 3.0 | Change | v2.0 | v3.0 ||---|---|---|| **Primary faithfulness** | Retraining-based AUDAC | **Frozen-model** masking-based AUDAC (no retraining) || **AUDAC values** | RF 0.679, XGB 0.662, LGB 0.688 | **RF 0.352, XGB 0.348, LGB 0.346, MLP 0.442, LSTM 0.463** || **Flip rate** | Malicious only | **Both malicious AND benign** (trees: 100%/0%) || **Per-instance SHAP-LIME ρ** | Negative (−0.18, buggy) | **Positive** (RF 0.72, XGB 0.43, LGB 0.52, MLP 0.58) || **LOSO** | RF only | **All 4 sklearn models** || **Feature ablation** | dst_port_id only | **4 features** (dst_port_id, src_port_class, hour_of_day, ip_reputation_score) || **Multi-seed** | RF only (3 seeds) | **All 6 models** × 3 seeds || **Service inventory** | 10 services listed | **2,309 unique ports** documented || **Transformer** | Negative control (F1 = 55.9%) | **Convergence failure** (F1 = 0.00 on 2/3 seeds) || **Labels** | "True labels" | **"Operational annotations"** || **Scripts** | 6 scripts | **10 scripts** || **Result files** | 2 CSVs | **7 CSVs + 2 TXT** | --- ## Overview This repository contains the complete reproducibility package for the manuscript *"Are Explanations Faithful Under Fire? Evaluating XAI Reliability on Live Honeypot Intrusion-Detection Telemetry."* The central contribution is a **frozen-model faithfulness evaluation** of model-agnostic explanations (SHAP, LIME) on live adversarial honeypot traffic. Each model is trained once, saved as a frozen checkpoint, and evaluated without retraining — directly measuring whether the trained model relies on the features SHAP identified. --- ## Repository Structure ```.├── README.md│├── # ── Primary Analysis: Frozen-Model Faithfulness ──├── frozen_model_faithfulness.py # Trains, saves, evaluates frozen models (6 models × 3 seeds)├── frozen_fixes.py # Corrects LIME name-matching bug + MLP SHAP Pipeline error├── frozen_faithfulness.csv # Primary results: 18 rows (6 models × 3 seeds)├── frozen_fixes_results.txt # Corrected SHAP-LIME ρ and MLP AUDAC│├── # ── Service Generalisation (Phase 4) ──├── phase4_generalisation.py # LOSO all models, multi-feature ablation, CIs├── loso_all_models.csv # LOSO F1 for RF, XGBoost, LightGBM, MLP (48 rows)├── multi_feature_ablation.csv # 4-feature ablation + combined removal (20 rows)├── per_service_all_models.csv # Per-service F1 for all 4 models (48 rows)├── service_inventory.csv # Complete inventory: 2,309 unique destination ports├── phase4_results.txt # Full Phase 4 output with bootstrap CIs│├── # ── Secondary Analysis: Retraining-Based ──├── run_faithfulness_retraining.py # Retraining-based AUDAC (secondary robustness analysis)├── xai_faithfulness_metrics.py # Faithfulness metrics library (AUDAC, flip rate, sparsity)├── faithfulness_retraining_results.csv # Retraining-based results (6 models, seed=42)│├── # ── Supporting Scripts ──├── per_service_evaluation.py # Per-service F1, LOSO, dst_port_id ablation (RF only, v2)├── bootstrap_ci.py # Bootstrap 95% CIs for AUDAC and flip rate├── reviewer_experiments.py # 6 reviewer-response analyses in one pass├── benchmark_diagnostics.py # Leakage audit and de-duplication analysis├── unsw_corrected_split_rerun.py # Corrected UNSW-NB15 official split re-run│├── # ── Benchmark Results ──└── verified_benchmark_results.csv # UNSW-NB15 and CIC-IDS2017 (6 models)``` --- ## Key Results ### 1. Primary Analysis: Frozen-Model Faithfulness (Table 9) Models trained once → saved as frozen checkpoints (.pkl/.keras) → top-k SHAP features masked with training-set medians → **no retraining**. This directly measures whether the trained model depends on the features SHAP identified. | Model | F1 | Frozen AUDAC ↓ | Mono. | Flip Mal | Flip Ben | Global ρ | Per-inst ρ ||---|---|---|---|---|---|---|---|| RF | 88.6% | **0.352** | Strict | 100% | 0% | 0.87 | 0.72 || XGBoost | 88.2% | **0.348** | Essential | 100% | 0% | 0.93 | 0.43 || LightGBM | 88.1% | **0.346** | Essential | 100% | 0% | 0.85 | 0.52 || MLP | 87.7% | **0.442** | Mostly | 98.8% | 2.2% | 0.87 | 0.58 || LSTM | 83.4% | **0.463** | Mostly | 87.0% | 13.6% | N/A | N/A || Transformer | 67.0% | 0.670 | N/A | 0% | N/A | N/A | N/A | **Monotonicity tiers:** Strict = no violations (k=0–10); Essential = violations only near chance; Mostly = one real violation, overall downward; Anti = accuracy rises (Transformer on retraining — excluded from primary). **Bootstrap 95% CIs (frozen, seed=42):**- AUDAC: RF [0.338, 0.366], XGBoost [0.332, 0.365], LightGBM [0.330, 0.362], MLP [0.425, 0.460]- Flip-rate: RF [1.000, 1.000], XGBoost [1.000, 1.000], LightGBM [1.000, 1.000], MLP [0.970, 1.000] **Key finding:** Tree ensembles flip 100% of malicious predictions but 0% of benign predictions when top SHAP features are masked — confirming that the identified features are specifically diagnostic of malicious behaviour. ### 2. Multi-Seed Variance (seeds 42, 7, 2024) | Model | F1 (mean ± SD) | AUDAC (mean ± SD) ||---|---|---|| RF | 0.891 ± 0.005 | 0.341 ± 0.018 || XGBoost | 0.884 ± 0.003 | 0.428 ± 0.086 || LightGBM | 0.888 ± 0.006 | 0.378 ± 0.027 || MLP | 0.881 ± 0.004 | 0.677 ± 0.024 || LSTM | 0.851 ± 0.014 | 0.461 ± 0.018 || Transformer | 0.223 ± 0.387 | 0.223 ± 0.387 | Transformer achieves F1 = 0.00 on two of three seeds — catastrophic convergence instability on tabular data. ### 3. SHAP-LIME Cross-Method Consistency (corrected in v3) | Model | Global ρ | Per-instance median ρ | IQR ||---|---|---|---|| RF | 0.87 | 0.72 | [0.63, 0.82] || XGBoost | 0.93 | 0.43 | [0.30, 0.59] || LightGBM | 0.85 | 0.52 | [0.38, 0.68] || MLP | 0.87 | 0.58 | [0.48, 0.68] | **Bug fix in v3:** Per-instance values in v2 were negative (median ρ = −0.18) due to a LIME feature-name matching error. LIME with `discretize_continuous=True` returns names like `"payload_length_bytes <= 245.50"` which did not match the original feature names during rank-correlation computation. Corrected in `frozen_fixes.py` using substring matching. ### 4. Service Generalisation (Phase 4 — new in v3) **Service inventory:** 2,309 unique dst_port_id values observed during the 21-day deployment:- 10 deliberately deployed honeypot services (ports 21, 22, 25, 53, 80, 443, 3306, 3389, 5900, 6379) — 651–744 sessions each- 2 additional ports with substantial traffic (8080 HTTP-alt, 8443 HTTPS-alt)- ~2,297 ephemeral scanning/probing ports (1–2 sessions each) **LOSO — all 4 sklearn models (12 services with ≥10 test samples):** | Model | Mean F1 Drop | Services Improved ||---|---|---|| RF | −0.10 pp | 4/12 || XGBoost | +0.54 pp | 4/12 || LightGBM | −0.07 pp | 6/12 || MLP | −0.43 pp | 7/12 | **Multi-feature ablation (all models):** | Feature Removed | RF | XGB | LGB | MLP ||---|---|---|---|---|| dst_port_id | −0.35 pp | −0.51 pp | −0.38 pp | −1.42 pp || src_port_class | +0.61 pp | +0.89 pp | +0.14 pp | +0.63 pp || hour_of_day | +0.81 pp | +1.51 pp | +0.38 pp | +0.96 pp || ip_reputation_score | −0.74 pp | +0.80 pp | +0.25 pp | −0.46 pp || **ALL FOUR** | **+2.80 pp** | **+3.36 pp** | **+2.68 pp** | **+5.43 pp** | Negative drop = F1 *improves* when that feature is removed. Removing dst_port_id improves F1 for all models — no shortcut learning. Removing all four service-related proxies costs only 2.7–5.4 pp, confirming that the core detection signal resides in behavioural features. ### 5. Secondary Analysis: Retraining-Based Faithfulness Retained from v2 as a robustness check. For each removal step k, the model is retrained on remaining features — allowing adaptation, producing higher (more conservative) AUDAC. | Model | F1 | AUDAC ↓ | Flip % | SHAP-LIME ρ ||---|---|---|---|---|| RF | 88.59% | 0.634 | 100% | 0.93 || XGBoost | 88.27% | 0.617 | 98.2% | 0.89 || LightGBM | 88.17% | 0.633 | 100% | 0.86 || MLP | 87.01% | 0.676 | 99.6% | 0.87 || LSTM | 83.15% | 0.630 | 12.8% | 0.93 || Transformer | 59.16% | 0.603 | 15.0% | 0.70 | ### 6. Benchmark Comparison (Table 10) | Model | UNSW-NB15 F1 | UNSW-NB15 AUC | CIC-IDS2017 F1 | CIC-IDS2017 AUC ||---|---|---|---|---|| XGBoost | 92.31% | 98.50% | 99.82% | 99.98% || RF | 92.24% | 98.29% | 99.80% | 99.98% || LightGBM | 92.06% | 98.49% | 99.90% | 99.99% || LSTM | 91.32% | 98.03% | 98.45% | 99.96% || MLP | 91.04% | 97.96% | 98.84% | 99.94% || Transformer | 90.77% | 96.86% | 97.27% | 99.92% | Note: different feature schemas (42 features for UNSW-NB15, 78 for CIC-IDS2017 vs 13 for the honeypot), different label definitions, and different preprocessing pipelines prevent direct comparison. Lower live-traffic scores may reflect adversarial difficulty but should not be interpreted as a causal relationship. --- ## Script Descriptions ### `frozen_model_faithfulness.py` — Primary Faithfulness (NEW in v3) The central experiment addressing the peer-review requirement for frozen-model evaluation. For each of 6 models × 3 seeds:1. Trains the model on the 85% training partition (n = 9,265)2. Saves a frozen checkpoint to `frozen_models/` (.pkl for sklearn, .keras for TensorFlow)3. Loads the frozen checkpoint for all subsequent evaluations4. Computes SHAP values (TreeExplainer for trees, KernelExplainer for deep models)5. Runs masking-based AUDAC (k = 0, 1, ..., 13) — **no retraining**6. Computes flip rate for **both** malicious and benign

View source

Similar papers

#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Open access Mar 2024

LLM-based agents for automating the enhancement of user story quality: An early report

The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.

Zheying Zhang, M. Rayhan, Tomas Herda et al. · 48 citations · ⚡4
#computer vision Review Mar 2024

System for systematic literature review using multiple AI agents: Concept and an empirical evaluation

This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.

Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al. · 44 citations · ⚡2
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations
#artificial intelligence Conference Open access Jun 2018

The Key Concepts of Ethics of Artificial Intelligence

It is suggested that the focus on finding keywords is the first step in guiding and providing direction for future research in the AI ethics field.

Ville Vakkuri, P. Abrahamsson · 39 citations · ⚡2

Related blog posts

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.