Benchmark-Shift-Aware Intrusion Detection for Evolving Network Traffic: Cross-Dataset Generalization, Calibrated Alerting, and Score-Orientation Diagnostics
Modern intrusion detection systems (IDSs) are often evaluated under matched training and test conditions, whereas deployment environments involve changing traffic distributions, heterogeneous feature-generation pipelines, and shifting attack prevalence. This study investigates benchmark-shift-aware intrusion detection through harmonized cross-dataset evaluation of HIKARI-2021, CICIDS2017, and a CICIoT2023 sample subset. Two payload-free feature spaces are constructed: Rich-64 for detailed HIKARI-2021/CICIDS2017 analysis and Minimal-13 for three-way comparison. Using XGBoost, a supervised Transformer, and a masked-feature self-supervised Transformer, we evaluate discrimination, calibration, threshold transfer, alert-budget behavior, chronological robustness, and score-orientation stability. Across five in-domain XGBoost settings, observed false-positive rates were 4.94–5.40%, and F1-scores ranged from 0.507 to 0.995. Under strict Rich-64 HIKARI-2021-to-CICIDS2017 transfer, all models had zero recall at source-derived thresholds, with two showing inverted score orientation. In the reverse direction, XGBoost reached an 18.9% target false-positive rate, while a nominal 5% target-side alert budget yielded F1 = 0.112. Chronological evaluation further showed that improved ranking metrics did not guarantee stable validation-derived operating behavior. The study provides a reproducible diagnostic framework for evaluating IDS robustness under evolving benchmark conditions.