Deployment-oriented evaluation of deep learning network intrusion detection systems
Abstract
Deep learning models are now used as a security layer in network intrusion detection systems. These models are often used in Internet of Things (IoT), industrial Internet of Things (IIoT), software-defined networking (SDN), and future internet communication environments. Most of the models reviewed in literature are still trained and tested on the same dataset under closed-set evaluation conditions; these models are also required to tolerate traffic distribution changes, unseen attacks, and uncertain predictions. Review articles from 2024 to 2026 focus on three main concerns, namely, cross-dataset generalization under domain shift, open-set intrusion recognition, and reliability through calibration or uncertainty-aware evaluation. Many of the reviewed models have been developed using isolated datasets and assumptions that weaken direct comparisons among them. This makes it difficult to judge deployment readiness across other environments. Most studies in literature address only one of these concerns instead of all three; across 29 empirical studies published between 2024 and 2026, none have evaluated all three concerns together. Several of these studies report high benchmark accuracies, but the reviewed corpus lacks joint evidence that the models can handle new attacks and new datasets while remaining reliable in IoT, IIoT, SDN, and future internet environments. Future works therefore require better cross-dataset testing, more evaluations of unseen attacks, and reliability checks in real-world secure communication environments.