Skip to content
Conference Open access

Semantic Feature Extraction from PE Headers for Malware Classification

2026 · International Conference on Security and Cryptography · pp. 1115-1120 · 0 citations · 17 references
Computer Science

TL;DR

A 60-feature taxonomy is presented that matches EMBER’s F1 in a directly compared head-to-head while using 39 × fewer features, a benign-source hold-out rules out a single-source artefact for the dominant feature, and Wilson-95% PPV bounds under deployment priors quantify what practitioners face at sub-1% malware prevalence.

Abstract

: Portable Executable (PE) malware classifiers are routinely benchmarked on malware-only data or against small homogeneous benign corpora. Recent surveys (Ucci et al., 2019; Aboaoja et al., 2022) note that most existing PE-feature studies select attributes by availability or precedent rather than by security rationale, and large benchmarks such as EMBER (Anderson and Roth, 2018) group 2 , 381 features only by extraction source while Ahmadi et al. (Ahmadi et al., 2016) similarly use > 1 , 800 features without semantic categorisation. We argue this practice misrepresents which signals a deployed detector actually relies on: the relative importance of the same 60 PE-header features changes substantially when benign samples are added to the evaluation, and shifts further as the benign corpus is diversified beyond a single source. To support this claim we organise 60 PE-header attributes into seven security-rationale categories (Structure Integrity, Execution Context, Memory Layout, Security Posture, Code Characteristics, Resource/Import, Anomaly Indicators) and evaluate on 1,263 MalwareBazaar samples plus 1,132 benign PE files (170 SysInternals + 962 DikeDataset (Iosif, 2021)). The 60-feature taxonomy matches EMBER’s F1 in a directly compared head-to-head while using 39 × fewer features, a benign-source hold-out rules out a single-source artefact for the dominant feature, and Wilson-95% PPV bounds under deployment priors quantify what practitioners face at sub-1% malware prevalence.

Read PDF

Similar papers

Open access Jun 2025

Enhancing android malware detection with retrieval-augmented generation

This work compiled a dataset of benign and malicious APKs and performed static analysis to extract features such as code structure, permissions, and manifest file content, without executing the apps, and used an LLM to generate high-level functional descriptions of APKs.

Saraga Sakthidharan, S. Anagha, Dincy R. Arikkat et al. · 2 citations
Conference Aug 2026

Disagreement-Aware Multi-View Stacking for Robust Malware Classification

Static malware family classification can use evidence from raw byte content, disassembly, and Portable Executable metadata. Each representation captures different characteristics of a malware sample. We propose a multi-view stacking framework for Microsoft BIG 2015 that represents each sample through seven static featu...

Kim-Huu Tran, Nguyen-Huy Le-Huu, Quoc-Huy Nguyen et al. · 0 citations
Open access 2026

MalBERT-Temporal: Transformer-Based Zero-Day Malware Detection in Windows Executable Binaries Under Strict Temporal Isolation

In current malware detection benchmarks, random train/test splits are commonly used, allowing temporal leakage to occur and obscuring performance degradation caused by concept drift. Moreover, traditional classifiers that operate on a one-dimensional PE feature vector do not explicitly model long-range interactions bet...

Manar Alanazi, Israa Alsiyat · 0 citations
Conference Aug 2026

Explainable Malware Detection from Noisy API Sequences with RAG-Based MITRE ATT&CK Mapping

As sophisticated evasion techniques like polymorphism and staged execution increasingly neutralize conventional signature-based defenses, dynamic API sequence analysis has emerged as an effective approach for malware detection. However, extracting actionable intelligence from noisy execution logs while maintaining mode...

Dat Quoc Phan, Tien Duc Anh Hao, Nghi Hoang Khoa et al. · 0 citations
Open access Sep 2026

SHAP-Guided Opcode Feature Selection for Lightweight and Explainable IoT Malware Detection — and What the Explanations Revealed About the Dataset

Opcode-frequency models detect IoT malware with near-perfect accuracy, yet typically with thousands of n-gram features and no account of which instructions drive a verdict. We propose a framework in which SHAP values are not a post-hoc report but the feature-selection criterion itself: a reference LightGBM model is exp...

Suad Shatti Azeez · 0 citations
Open access 2026

Cross-Dataset Feature Discriminability in Static Malware Detection: A CVFR-Based Study Across Four Heterogeneous Benchmarks

Static malware detection with machine learning relies critically on which features are extracted from Portable Executable (PE) binaries. Most prior work evaluates feature selection on a single benchmark, leaving open whether features generalize across heterogeneous datasets. We investigate cross-dataset feature discrim...

D. Truong · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.