MalBERT-Temporal: Transformer-Based Zero-Day Malware Detection in Windows Executable Binaries Under Strict Temporal Isolation
Abstract
In current malware detection benchmarks, random train/test splits are commonly used, allowing temporal leakage to occur and obscuring performance degradation caused by concept drift. Moreover, traditional classifiers that operate on a one-dimensional PE feature vector do not explicitly model long-range interactions between structurally distant feature groups. This study introduces a Transformer-based malware detection approach named MalBERT-Temporal that reshapes the 2,381-dimensional BODMAS PE feature vector into 16 contiguous feature-group tokens and then processes these tokens with Transformer encoder layers employing multi-head self-attention to model interactions among all tokens. The proposed approach is evaluated using a strict temporal protocol in which the training data comprise only pre-2020 samples, and the test data span the complete 2020 evaluation timeline. Each model is independently calibrated using a fixed validation set with a false-positive-rate budget of 0.1% and is evaluated monthly and on malware families unseen during training. At the calibrated security threshold, MalBERT-Temporal achieves an F1-score of 97.84% with a false-positive rate of 0.138%, outperforming a 1-D CNN baseline, which achieves an F1-score of 92.01% and a false-positive rate of 3.388%, and a Random Forest baseline, which achieves an F1-score of 67.29%. Welch’s t-tests confirm statistically significant differences in monthly F1-scores between the proposed approach and both baselines. Moreover, MalBERT-Temporal maintains monthly F1-scores within a 2.2-point band throughout the nine-month evaluation period, indicating improved robustness to temporal distribution shift. A component-wise ablation traces the improvement to the tokenized self-attention mechanism itself: substituting a token-wise feed-forward block for self-attention while keeping every other component costs 2.73 F1 points, and removing the tokenization costs 1.85 points, while positional embeddings, the CLS token, multi-scale pooling, encoder depth, and nonlinearity are each worth 0.47 points or less. A matched random-split control that leaves the model, the preprocessing, the training budget, and the calibration procedure unchanged gives an F1-score of 98.86%, which confirms that chronological evaluation is the stricter protocol.