Skip to content
Open access

MalBERT-Temporal: Transformer-Based Zero-Day Malware Detection in Windows Executable Binaries Under Strict Temporal Isolation

2026 · International Journal of Advanced Computer Science and Applications · 0 citations · 29 references

Abstract

In current malware detection benchmarks, random train/test splits are commonly used, allowing temporal leakage to occur and obscuring performance degradation caused by concept drift. Moreover, traditional classifiers that operate on a one-dimensional PE feature vector do not explicitly model long-range interactions between structurally distant feature groups. This study introduces a Transformer-based malware detection approach named MalBERT-Temporal that reshapes the 2,381-dimensional BODMAS PE feature vector into 16 contiguous feature-group tokens and then processes these tokens with Transformer encoder layers employing multi-head self-attention to model interactions among all tokens. The proposed approach is evaluated using a strict temporal protocol in which the training data comprise only pre-2020 samples, and the test data span the complete 2020 evaluation timeline. Each model is independently calibrated using a fixed validation set with a false-positive-rate budget of 0.1% and is evaluated monthly and on malware families unseen during training. At the calibrated security threshold, MalBERT-Temporal achieves an F1-score of 97.84% with a false-positive rate of 0.138%, outperforming a 1-D CNN baseline, which achieves an F1-score of 92.01% and a false-positive rate of 3.388%, and a Random Forest baseline, which achieves an F1-score of 67.29%. Welch’s t-tests confirm statistically significant differences in monthly F1-scores between the proposed approach and both baselines. Moreover, MalBERT-Temporal maintains monthly F1-scores within a 2.2-point band throughout the nine-month evaluation period, indicating improved robustness to temporal distribution shift. A component-wise ablation traces the improvement to the tokenized self-attention mechanism itself: substituting a token-wise feed-forward block for self-attention while keeping every other component costs 2.73 F1 points, and removing the tokenization costs 1.85 points, while positional embeddings, the CLS token, multi-scale pooling, encoder depth, and nonlinearity are each worth 0.47 points or less. A matched random-split control that leaves the model, the preprocessing, the training budget, and the calibration procedure unchanged gives an F1-score of 98.86%, which confirms that chronological evaluation is the stricter protocol.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.