Skip to content
Open access

Cross-Dataset Feature Discriminability in Static Malware Detection: A CVFR-Based Study Across Four Heterogeneous Benchmarks

2026 · IEEE Access · Vol 14, pp. 115424-115439 · 0 citations · 27 references
Computer Science

Abstract

Static malware detection with machine learning relies critically on which features are extracted from Portable Executable (PE) binaries. Most prior work evaluates feature selection on a single benchmark, leaving open whether features generalize across heterogeneous datasets. We investigate cross-dataset feature discriminability by quantifying the consistency of important features across heterogeneous malware datasets. We introduce Cross-Validated Feature Ranking (CVFR), a lightweight wrapper that aggregates Random Forest (RF) feature importance across repeated cross-validation runs to attach per-feature confidence intervals and a stability score (SNR) to an otherwise standard ranking. CVFR is not a new feature-selection algorithm and is not claimed to improve accuracy; it is positioned as the uncertainty-quantification instrument that enables the cross-dataset discriminability study. RF, Multilayer Perceptron (MLP), LightGBM, and One-Class SVM are evaluated under a two-level stratified cross-validation protocol (25 repeated evaluations) with within-fold SMOTE on four PE-header benchmark datasets: EMBER, BIG15, Malicia, and BODMAS. RF achieves the highest F1 and AUC-ROC on every dataset (mean 98.72% F1, 0.998 AUC), consistently outperforming MLP across all 25 fold-seed pairs (Wilcoxon signed-rank test, $p \lt 0.0001$ ), and remaining best or tied against a competitive LightGBM gradient-boosting baseline. An ablation study confirms that CVFR achieves predictive performance comparable to single-run impurity ranking and permutation importance, while additionally providing uncertainty estimates for feature importance at lower computational cost than permutation importance. Jaccard analysis reveals limited feature overlap (7.9–27.6%) among independently collected datasets, while datasets sharing the same extraction schema (EMBER and BODMAS) reach 83.8% overlap. Transfer experiments show that the high-overlap pair (EMBER $\leftrightarrow $ BODMAS, 83.8% Jaccard) experiences larger F1 degradation (20–36 pp) than the low-overlap pair (BIG $15\leftrightarrow $ Malicia, 8–13 pp), suggesting that feature-space overlap alone is insufficient to predict transferability and that temporal distributional shift may exert a stronger influence on cross-dataset generalization. Because the transfer analysis covers only two dataset pairs (two seeds each) that co-vary simultaneously in feature overlap, temporal collection gap, and class balance, this transfer observation is hypothesis-generating rather than statistically conclusive. Subject to that caveat, the findings indicate that feature importance should be validated separately for each target dataset and that successful cross-dataset deployment is likely to require explicit domain adaptation, even when datasets share the same feature-extraction schema.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.