XAI-Enabled Intelligent Detection and Mitigation of Zero-Day Attacks in Healthcare IoT
Abstract
IoMT presents a life-critical healthcare infrastructure to an ever-changing threat terrain, where the worst attacks are by definition the ones that have never been seen. Conventional detection souls are trained and tested on the same shackle distribution, reporting growth accuracies which to enwrinkle inside of rise (zero-day) traffic. This paper proposes an explainable, cost-sensitive stacking framework that is explicitly tested under a cross-dataset zero-day protocol: the detector is trained purely on CICIoMT2024's healthcare-IoMT threat landscape including 21 attack families absent in training from full network traffic captures of similar timestamped IoT devices running second daily command & control patterns of attacks labelled as seen/unseen using human interpretation rather than naive string matching. This framework combines four supervised base learners (a multilayer perceptron, a one-dimensional convolutional network, a TabNet-style attentive network and gradient-boosted trees) with the anomaly score generated by a benign-only autoencoder as an additional meta-feature which allows the model to mark attacks it has never been trained on. Cost-sensitive weighting is used for class imbalance and all hyperparameters are tuned using Bayesian (Tree-structured Parzen Estimator) optimization.The tuned model reaches macro-F1 = 0.758, Matthews correlation coefficient (MCC) = 0.523, ROC-AUC = 0.968 and PR-AUC = 0.998 on the zero-day set, with pooled recall on the unseen families of 21 is high at 0.986 statistically indistinguishable from its recall of semantically-seen families evaluated cross-dataset (recall=0.985), and outperforms the strongest ablation baseline significantly (McNemar chi-square= 108.9, p <10-24). Tree- and Gradient-Based Explainers Identifying a Compact Set of Flow-Statistical Drivers: A SHAP audit combines tree- and gradient-based explainers which instead identify five features (rst_count, Number, IAT, Weight, Magnitude) that concentrate the zero-day signal into a small set of flow-statistical drivers. We report two genuine limitations (and no fake ones): recall on stealthy application-layer attacks (injection, brute-force, single-request exploits) is only 0.569 because their signature lives in payload rather than flow statistics; and the benign false-positive rate increases from 0.033 in-distribution to 0.534 cross-dataset correctable via retraining but not by nefarious rutin disclosure which otherwise affect all seven models we evaluated