MalXGraph-ViT: A Hybrid Transformer-graph Neural Network Framework with Explainable AI for Robust Malware Detection and Classification on Emerging Public Benchmarks
Abstract
The exponential growth of malware variants, propelled by automated obfuscation, polymorphism, and AIassisted code generation, has rendered signature-based and heuristic detection ineffective.This paper proposes MalXGraph-ViT, a multi-stage framework integrating grayscale visual feature extraction, Global Context Vision Transformers (GCViT), Graph Neural Networks over opcode-import dependency graphs, and Explainable Artificial Intelligence (XAI) via SHAP and Grad-CAM.From the EMBER2024 release we use 1,040,000 labelled Windows-PE samples and its separate 50,000-sample evasive-challenge partition.WinMalware2025 (401,508 samples, 16 classes, 160 families) and BODMAS (57,293 malware + 77,142 benign, 581 families) are evaluated together with Malimg, MaleVis, and CIC-MalMem-2022.The pipeline attains 99.62% binary accuracy on the held-out EMBER2024-Win64 test partition, 98.97% macro-F1 on the BODMAS top-20-family subset, and AUROC = 0.991 on the EMBER2024 evasive split, ahead of GCViT (0.971) and B_ViT (0.969) retrained by us on identical splits.SHAP and Grad-CAM evidence converges on the same opcode-import regions, giving a mean Explainability-Agreement Score (EAS) of 0.71.Unlike graph-only or vision-only detectors, MalXGraph-ViT couples byte-image tokens with individual graph-node tokens through cross-modal attention; removing the graph branch alone costs 2.0 pp of evasive AUROC.A calibrated Bayesian few-shot head reaches 72.9% 5-shot zero-day accuracy on EMBER2018 (ECE = 2.2%), +9.5 pp above the BayesMAML-E baseline (63.4%).Code, split manifests, seeds, per-seed metrics, and baseline provenance are public at https://github.com