Cross-Dataset Generalization Framework for Cybercrime Detection Using CIC-IDS2017, CIC-Phishing2019, and Malicious Uniform Resource Locator (URL) Data
Abstract
—Proposed cybercrime detection models demonstrate satisfactory performance on specific benchmark datasets; however, they are not always robust across diverse, practical scenarios. This paper examines the generalization performance of a cross-dataset measure for cybercrime detection across network traffic, phishing email, and malicious Uniform Resource Locator (URL) domains. The work combines three popular security benchmarks Canadian Institute for Cybersecurity Intrusion Detection System Dataset2017 (CIC-IDS2017), Canadian Institute for Cybersecurity Phishing Dataset2019 (CIC-Phishing2019), and Malicious URL 2020 in a single preprocessing and learning pipeline to avoid dataset-specific bias. One or more datasets are used to train models, which are then directly evaluated on previously unseen datasets to verify transferability across distribution shifts. We will discuss performance while considering accuracy, F1 − Score, Receiver Operating Characteristic (ROC), and fold-to-fold stability. Experiments demonstrate that direct training with a single random source results in significant performance deterioration, with a 23% decrease in F1 − Score when applied to unseen datasets. In contrast, the degradation caused by the proposed framework is kept below 10% and maintains Receiver Operating Characteristic–Area Under the Curve (ROC–AUC) values consistently above 0.90. Paired significance testing demonstrates that the gains in robustness are highly significant ( p < 0.01). The results indicate that cross-dataset evaluation is essential for the