This study proposes a novel hybrid malware representation framework that integrates Convolutional Autoencoder (CAE)-based latent structural learning with Local Binary Pattern (LBP)-based texture feature extraction for robust malware classification and represents one of the first comprehensive investigations of hybrid latent-texture representation learning within a memory-forensics malware visualization setting.
Abstract
Visualization-based malware detection has recently gained significant attention for binary and multiclass malware classification using machine learning and deep learning techniques. However, existing visualization-based frameworks still face several important limitations, including insufficient robustness evaluation, limited cross-dataset validation, restricted malware diversity, and difficulty distinguishing visually similar and noise-sensitive malware families. In many cases, the visual similarity between malware classes and the presence of perturbations negatively affect feature extraction quality, leading to degraded classification performance and reduced generalization capability. To address these challenges, this study proposes a novel hybrid malware representation framework that integrates Convolutional Autoencoder (CAE)-based latent structural learning with Local Binary Pattern (LBP)-based texture feature extraction for robust malware classification. To the best of our knowledge, this study represents one of the first comprehensive investigations of hybrid latent-texture representation learning within a memory-forensics malware visualization setting while jointly addressing robustness, perturbation resilience, scalability, and cross-dataset generalization through a unified evaluation framework. The proposed framework combines global hierarchical representations learned through CAE with fine-grained local texture descriptors extracted using LBP to improve the discrimination of visually similar malware families and enhance robustness against noisy visualization conditions. The extracted features are subsequently evaluated using multiple machine learning classifiers, where XGBoost achieved the highest performance with an accuracy of 99.90%, precision of 99.79%, recall of 99.92%, and F1-score of 99.85%. To comprehensively evaluate the proposed framework, extensive experiments are conducted using both a memory-forensics malware dataset and the large-scale BODMAS dataset containing 134,435 PE malware samples spanning 581 malware families. The experimental evaluation incorporates cross-validation, ablation analysis, robustness assessment under multiple perturbation conditions, and feature-space visualization analysis. The results demonstrate that the proposed CAE+LBP framework consistently outperforms standalone feature extraction approaches and conventional end-to-end CNN models while maintaining strong robustness and cross-dataset generalization capability across diverse malware distributions and noisy conditions.
A Hybrid Neural Network–Convolutional Neural Network (NN–CNN) Deep Learning Framework for malware detection, malware-family classification, and malware-variant identification and considers two important issues in practical malware detection: model explainability and generalization to previously unseen malware.
Chioma Grace Nwankwo, B. C. Amanze, Ikechukwu Amaefule· World Journal of Advanced Re...· 0 citations
Malware is a serious threat in the cybersecurity area because of its dynamic nature, the variety of malware families, stealth, propagation and the capability of evading traditional security products. Therefore, proper malware detection and classification are crucial for detecting malicious software and for securing computer systems from unauthorized access and data stealing, and for disrupting systems. This study covers all the bases when it comes to deep learning approaches for malware detection and classification. It covers the principles, different forms of malware, how to detect deep learning malware, how to represent data, obtaining features, and applications. The traditional detection methods are described with their drawbacks, namely based on signature, behavioral and heuristic methods. The report also delves into the methodologies used by deep learning to classify malware, namely CNNs and Bidirectional Long Short-Term Memory (BiLSTM) networks. BiLSTM models excel at learning sequential features from code-or behavior-related data, whereas CNN-based representation learning approaches excel at learning spatial features from malware representations. Moreover, the various detection techniques (static, dynamic and hybrid) are discussed so that their role in malware analysis can be understood. The survey identifies the current challenges and gaps in research and emphasizes the need for strong, scalable and adaptive deep-learning models to combat new malware threats and enhance cybersecurity protection.
Shivani Jain· International Journal of Cyb...· 0 citations
It is indicated that integrating static and dynamic features at the representation level can improve robustness and classification performance in image-based malware detection, particularly under variations in feature availability.
Anis Elgarduh, A. Zainal, Fuad A. Ghaleb et al.· IEEE Access· 0 citations
This paper introduces a compact convolutional neural network (CNN) architecture integrated with a multi-head attention mechanism to enhance feature discrimination in malware family classification. The novelty lies in combining attention-based refinement within a lightweight framework that achieves comparable accuracy to larger model while maintaining computational efficiency. The integrated multi-head attention module emphasizes small yet highly discriminative regions, which often correspond to obfuscated or subtly modified code segments that traditional models overlook. In contrast to reverse engineering or handcrafted feature methods, the framework is fully automated and computationally efficient, requiring only 2.03 M parameters approximately one-fifth of the baseline CNN while reducing training time by 42% and inference latency by 37% on the same hardware configuration. Experiments were conducted on the widely used Malimg dataset, which comprises 9,342 grayscale images across 25 malware families, under multiple train–test splits (50/50 to 90/10). The results consistently demonstrate superior performance relative to the baseline CNN, the model achieving an average accuracy of 98.7 ± 0.3% across five randomized train–test splits, with a peak performance of 99% on the 80/20 partition. Data partitions were strictly disjoint to prevent leakage between training and test sets. Furthermore, the proposed model attains high weighted Precision, Recall, and F1 scores (≈0.99) and enhanced macro-averaged performance (up to 0.97), particularly enhancing classification for minority families that the baseline misclassified or failed to detect. Confusion matrix analysis further highlights that residual misclassifications occur primarily in families with limited samples or high morphological similarity. Comparative evaluation against recent studies confirms the superiority of the proposed approach in terms of both accuracy and computational efficiency. Despite these strengths, limitations remain, including dependence on labeled datasets, restricted interpretability, and the need for validation on more diverse, real-world malware corpora.
Mohammed Kasem Al-Bayati· Zanco Journal of Pure and Ap...· 0 citations
A systematic framework to enhance adversarial robustness is proposed, validated on the Malimg dataset and supersedes previous approaches by 13.15% in terms of the evasion rate and 37.34% in terms of retraining success.
Muhammad Arham Tariq, Allah Bux Sargano, Z. Habib et al.· International Journal of Inf...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.