Feature-Driven Framework for Interpretable Detection and Analysis of Android Malware Through Network and Application Behaviors
Abstract
This paper presents a complete framework for feature-driven Android malware detection that combines predictive modeling with explainable artificial intelligence to improve classification accuracy, interpretability, and operational dependability. The system initiates by examining network-flow metrics, including Source Port (SP), Initial Window Bytes Forward (IW), packet rates, and user activity indicators, derived from a publicly accessible dataset including 355,630 cases across four classifications: Adware, Scareware, SMS malware, and benign. Feature preprocessing utilizes the Variance Inflation Factor (VIF) analysis to eliminate multicollinearity to discern the most important variables, hence assuring a non-redundant and highly relevant feature collection. Subsequently, Gradient Boosting Classifier (XGBC) and Decision Tree Classifier (DTC) are augmented with metaheuristic optimization algorithms, Tailor Optimization Algorithm (TOA) and Spotted Deer Optimization Algorithm (SDOA), to boost convergence, stability, and generalization. Model interpretability is attained by Delta-XAI, which identifies SP, IW, and User Interaction/Write operations as the principal features, collectively representing approximately 80% of the explanatory power of the top features. Assessment by five-fold cross-validation indicates that the XGBC+TOA (XGTO) configuration attains an overall accuracy of 98.0%, exhibiting excellent precision and Matthews Correlation Coefficient (MCC), surpassing all other models. This framework combines feature-centric selection, enhanced learning, and explainable AI to deliver a dependable, interpretable, and pragmatic approach for identifying Android malware, thereby connecting data-driven analysis with practical cybersecurity implementations.