Skip to content
Open access

SHAP-Guided Opcode Feature Selection for Lightweight and Explainable IoT Malware Detection — and What the Explanations Revealed About the Dataset

Sep 2026 · Al-Noor Journal of Engineering Management and Computer Science · 0 citations · 6 references

Abstract

Opcode-frequency models detect IoT malware with near-perfect accuracy, yet typically with thousands of n-gram features and no account of which instructions drive a verdict. We propose a framework in which SHAP values are not a post-hoc report but the feature-selection criterion itself: a reference LightGBM model is explained fold-by-fold, opcodes are ranked by mean absolute SHAP contribution, and the smallest top-k set that preserves performance becomes the deployed model's input space. On the public CyberScienceLab ARM-32 opcode corpus (268 benign, 244 malicious samples), five selected features out of roughly 4,700 retain an F1 of 0.9939 ± 0.0035, statistically indistinguishable from the strongest of six baselines (Wilcoxon p = 0.121) under repeated stratified 10-fold cross-validation with selection performed inside each fold. The explanations also exposed a problem the accuracy tables hide. The single most influential token, an uppercase RET that is not an ARM-32 mnemonic, appears in 243 of 244 malware files and in no benign file — a collection artifact that may inflate reported accuracies on this dataset. We therefore introduce an artifact-resistant protocol that restricts the vocabulary to opcodes present in at least 5% of training files of each class, computed within each fold. Under this stricter regime the five-feature model reaches an F1 of 0.9913 ± 0.0051, trading 0.6 points of F1 against the best baseline (p = 0.046) for a 4x smaller feature set and explanations built into the pipeline. Interpretability is not an accessory: it is the mechanism that compressed the model and audited the data.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.