Machine Learning-Based Domain Risk Assessment for Cybersecurity Monitoring in Industrial Systems
Abstract
In the field of cybersecurity, malicious website classification plays a crucial role in protecting industrial systems. For this reason, research has been undertaken to analyze cybersecurity threats, with the long-term objective of developing methods for the effective detection and classification of malicious websites. This article evaluates the use of domain features for classifying malicious websites with machine learning methods. The feature vector consisted of 44 infrastructural, lexical, structural, and reputation-related characteristics. The model comparison included Logistic Regression, Support Vector Machine, Random Forest, AdaBoost, and XGBoost. Experiments were conducted on 247,730 URLs from the malicious_phish dataset (2021), with features extracted as part of this study in April 2026. The most predictive features were related to domain registration history, DNS infrastructure, and reputation-based rankings, as confirmed by ANOVA F-test, SHAP values, XGBoost gain, and permutation importance. Validation of the best-performing model on 1000 active phishing domains from the PhishDestroy list dated 30 May 2026, achieved a recall of 70%, while the application of a three-tier risk scale allowed 84.9% of domains to be flagged as malicious or suspicious.