Interpretable AI-Enhanced Machine Learning for Intelligent Spambot and Fake Followers Identification
TL;DR
The experimental results indicate that feature selection combined with CatBoost provides an effective approach for spambot prediction while explainability methods can provide additional insight into model decisions.
Abstract
Social spambots are automated accounts that imitate normal users and can be used to generate misleading, manipulative, or unwanted activity on social networking platforms. Accurate identification of such accounts is challenging because bot and human accounts may share similar profile and behavioural characteristics. This work presents a machine-learning based framework for spambot prediction using account-level and behavioural features. Genuine-user and social-spambot datasets were merged and assigned Human and Bot labels. Missing values were replaced with zero, labels were encoded numerically, features were shuffled and normalized using Min-Max scaling, and the resulting dataset was divided into 80% training and 20% testing subsets. Several machine-learning classifiers, including Random Forest, SVM, Decision Tree, XGBoost, LightGBM, Logistic Regression, Extra Trees, Naive Bayes, and AdaBoost, were evaluated using accuracy, precision, recall, and F1-score. As an extension, Chi-Square SelectKBest feature selection was used to retain 20 features before training a CatBoost classifier. The extended CatBoost model achieved 99.567% accuracy, precision, recall, and F1-score on the test set. To improve model interpretability, LIME and SHAP were incorporated to examine the contribution of input features to predictions. The experimental results indicate that feature selection combined with CatBoost provides an effective approach for spambot prediction while explainability methods can provide additional insight into model decisions.