Towards a comprehensive landscape of AI and generative AI in cybersecurity
Abstract
Conventional Intrusion Detection Systems (IDS) mechanisms based on signatures and anomalies struggle and have shown their limitations with novel and dynamic attack patterns, and it has become essential to discover the emerging potential offered by AI and GenAI. These models open new perspectives but also present significant risks that must be properly managed. This paper presents an in-depth comprehensive Systematic Mapping Study (SMS) of 110 relevant articles published between 2015 and 2025 related to AI/GenAI-based IDS. The proposed work offers a novel and integrated, comprehensive mapping of both defensive and offensive dimensions of AI/GenAI-enabled cybersecurity, a classification of AI models and Large Language Models (LLMs), a deep dive analysis of traditional and next generation cyber-attacks, synthesis datasets and methodologies utilized for evaluating AI-based Intrusion Detection System approaches. The findings highlight a substantial intensification in research efforts beginning in 2022. Regarding traditional Cyber Attacks (CA), malicious actors are taking advantage of AI/GenAI techniques to enhance existing conventional CA. In the meantime, academic research continues to focus on traditional attack types while leveraging the new capabilities offered by AI and GenAI. About studies on AI models that deal with IDS, we have noted that most articles are based on a hybrid approach combining different models. They are based on ensemble learning and AI boosting techniques, coupled with metaheuristic optimization algorithms for the implementation of the IDS pipelines. While supervised learning is still the dominant approach, semi-supervised, unsupervised, graph-based, and reinforcement learning also hold promise in detecting unknown and zero-day threats. In addition, we presented the evolution of LLMs timeline through several waves, from Transformer Architectures Based LLMs era to Hybrid multimodal models. We presented also the classification of the most used LLMs in cybersecurity field. Regarding the datasets for training and evaluating AI models related to IDS: NSL-KDD and UNSW-NB15 datasets have emerged as the primary benchmarks favoured by the scientific community.