Self-refining structures in large-language generative AI models for the generation of cybersecurity detection capabilities using machine-based and natural languages
Oct 2026· Zenodo (CERN European Organization for Nuclear Research)
Network Security and Intrusion Detection
Abstract
The cybersecurity threat landscape evolves rapidly placing increasing pressure on detectionapproaches that rely on static rules, manually curated signatures, or models trained on fixeddatasets. While modern security environments collect large volumes of telemetry data, theprocesses used to generate and maintain detection knowledge remain heavily dependent onongoing human effort. These approaches struggle to adapt as attack behaviours change. Thisthesis investigates whether autonomous generative systems can be used to generate andimprove cybersecurity detection knowledge without continuous human intervention. Ratherthan treating detection logic as a static artefact, the research examines how generative systemscan refine their behaviour over time through feedback-driven mechanisms. An artefact-drivenmethodology is adopted to design, implement, and evaluate two distinct but related generativesystems.The first system is a grammar-grounded generator that produces structured, machine-readableSQL detection queries directly from formal grammars. Generation is performed withoutexternal input, and syntactic correctness is guaranteed through grammar-constrained traversal.This system is extended by exposing internal decision points and introducing heuristic andweighted mechanisms that influence traversal behaviour based on prior outcomes. The secondsystem is a large language model–based pipeline designed to generate natural-languagecybersecurity reasoning and synthetic datasets. In this pipeline, generated outputs areautomatically evaluated using task-specific criteria, and feedback is used to guide regenerationuntil defined quality thresholds are met.The weighted generator was further evaluated in a controlled cybersecurity detectionexperiment using CIC-IDS2017-derived network-flow records. The selected SQL ruleachieved an F1 score of 0.7361 and balanced accuracy of 0.7389 on a held-out test set,compared with an F1 score of 0.6667 and balanced accuracy of 0.5000 for an always-maliciousbaseline. The LLM Judge was separately validated against blinded human ratings across 30generated responses. The results showed partial agreement, indicating that the Judge cansupport automated screening and refinement but should not be treated as a substitute for humanexpertise.By analysing these systems together, the thesis demonstrates that self-refinement can be treatedas a shared design principle across both structured and probabilistic generative architectures.Although the mechanisms differ, weighted traversal versus automated evaluation, both systemsexhibit outcome-driven adaptation that improves output quality over time. The findings showthat autonomous self-refinement can reduce reliance on static rules, fixed datasets, and manualcuration, providing a practical foundation for adaptive cybersecurity detection knowledgegeneration.
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.
P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al.· IEEE International Conferenc...· 110 citations· ⚡7
The findings show that speed related agile practices are used to a greater extent in comparison to quality practices, and that software startups who adopt the Lean Startup approach do not sacrifice quality for speed more than other startups do.
Jevgenija Pantiuchina, Marco Mondini, Dron Khanna et al.· International Conference on...· 84 citations· ⚡4
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 8, 2026
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.