Cross-Dataset Feature Discriminability in Static Malware Detection: A CVFR-Based Study Across Four Heterogeneous Benchmarks
Abstract
Static malware detection with machine learning relies critically on which features are extracted from Portable Executable (PE) binaries. Most prior work evaluates feature selection on a single benchmark, leaving open whether features generalize across heterogeneous datasets. We investigate cross-dataset feature discriminability by quantifying the consistency of important features across heterogeneous malware datasets. We introduce Cross-Validated Feature Ranking (CVFR), a lightweight wrapper that aggregates Random Forest (RF) feature importance across repeated cross-validation runs to attach per-feature confidence intervals and a stability score (SNR) to an otherwise standard ranking. CVFR is not a new feature-selection algorithm and is not claimed to improve accuracy; it is positioned as the uncertainty-quantification instrument that enables the cross-dataset discriminability study. RF, Multilayer Perceptron (MLP), LightGBM, and One-Class SVM are evaluated under a two-level stratified cross-validation protocol (25 repeated evaluations) with within-fold SMOTE on four PE-header benchmark datasets: EMBER, BIG15, Malicia, and BODMAS. RF achieves the highest F1 and AUC-ROC on every dataset (mean 98.72% F1, 0.998 AUC), consistently outperforming MLP across all 25 fold-seed pairs (Wilcoxon signed-rank test, $p \lt 0.0001$ ), and remaining best or tied against a competitive LightGBM gradient-boosting baseline. An ablation study confirms that CVFR achieves predictive performance comparable to single-run impurity ranking and permutation importance, while additionally providing uncertainty estimates for feature importance at lower computational cost than permutation importance. Jaccard analysis reveals limited feature overlap (7.9–27.6%) among independently collected datasets, while datasets sharing the same extraction schema (EMBER and BODMAS) reach 83.8% overlap. Transfer experiments show that the high-overlap pair (EMBER $\leftrightarrow $ BODMAS, 83.8% Jaccard) experiences larger F1 degradation (20–36 pp) than the low-overlap pair (BIG $15\leftrightarrow $ Malicia, 8–13 pp), suggesting that feature-space overlap alone is insufficient to predict transferability and that temporal distributional shift may exert a stronger influence on cross-dataset generalization. Because the transfer analysis covers only two dataset pairs (two seeds each) that co-vary simultaneously in feature overlap, temporal collection gap, and class balance, this transfer observation is hypothesis-generating rather than statistically conclusive. Subject to that caveat, the findings indicate that feature importance should be validated separately for each target dataset and that successful cross-dataset deployment is likely to require explicit domain adaptation, even when datasets share the same feature-extraction schema.