Skip to content
Review Open access

From static tasks to dynamic reasoning: a characterization framework and study of large language models in next-generation cybersecurity automation

Aug 2026 · International Journal of Information Security · Vol 25 · 0 citations · 88 references

TL;DR

This work surveys recent LLM-based systems across seven core domains and identifies the need for privacy-aware deployment, timely retrieval and knowledge maintenance for emerging threats, process-level evaluation tied to measurable security outcomes, and human oversight within controlled and hybrid automation workflows.

Abstract

Large Language Models (LLMs) are increasingly applied in cybersecurity, but most existing industry use cases focus on static, one-shot tasks such as classification, entity extraction, or summarization. While effective in narrow contexts, these applications fail to capture the complexity of real-world cybersecurity workflows, which often unfold over time, involve evolving inputs, and require multi-step reasoning. In this paper, we shift the focus toward dynamic cyber tasks—problems that demand context awareness, tool interaction, and adaptive decision-making. Our main goal is to define, analyze, and investigate the role of LLMs in automating these dynamic tasks. To achieve this, we introduce a characterization framework that profiles dynamic cyber tasks along four complementary dimensions: operational goal, knowledge grounding, collaboration mode, and cognitive complexity. We survey recent LLM-based systems across seven core domains: threat intelligence, data privacy and security, vulnerability detection, malware detection, intrusion detection, incident response and red teaming automation. Our analysis shows that current systems remain limited by privacy and deployment constraints, stale or incomplete threat knowledge, weak validation of feedback-driven actions, and insufficient evidence of operational benefit. We identify the need for privacy-aware deployment, timely retrieval and knowledge maintenance for emerging threats, process-level evaluation tied to measurable security outcomes, and human oversight within controlled and hybrid automation workflows. These findings clarify where LLMs can provide practical value and where conventional or hybrid approaches may remain more suitable.

Read PDF

Similar papers

Jul 2026

GARAGE: Characterizing the Automation Boundary in LLM-based Attack Graph Generation

GARAGE is introduced, a RAG-powered framework that converts fragmented CTI into an actionable, domain-specific knowledge base for automated attack graph generation and position GARAGE as a scalable TARA support tool within human-in-the-loop workflows, offering a comprehensive cost-performance analysis to guide its deployment across various LLM tiers.

Daekwon Pi, Sangho Lee, Young-Hun Lee et al. · 0 citations
Open access Jul 2026

Tool-Flow Taint Analysis for Data Exfiltration Defense in Large Language Model Agents

A comprehensive framework based on Tool-Flow Taint Analysis designed to mitigate data exfiltration in Large Language Model agents is introduced, providing a critical foundation for securing next-generation autonomous agents against sophisticated data-stealing attacks in enterprise environments.

Chun Tian, Hiu-Tung Li, Michelle Yu · 0 citations
Book Open access Aug 2026

OmniVul: A Holistic, Multi-Turn Conversational Benchmark for LLM-Based Vulnerability Assessment

With more than 20,000 Common Vulnerabilities and Exposures (CVEs) reported annually, software vulnerabilities represent a critical cybersecurity challenge. This volume has intensified the demand for automated detection and analysis, motivating the integration of large language models (LLMs) for such tasks. However, existing vulnerability benchmarks are not suitable for evaluating LLMs' capabilities in vulnerability assessment, as most of them 1) rely on narrow data sources, 2) lack deep context, and 3) focus on single-turn Q&A rather than realistic, multi-stage analyst workflows. To address this gap, we introduce OmniVul, a comprehensive multi-turn benchmark for LLM-based vulnerability assessment. OmniVul comprises 2,000 CVEs with question–answer pairs spanning 23 attributes, including detection, code localization, root cause analysis, and patch suggestion. We employ an automated workflow to aggregate multi-source data via Retrieval-Augmented Generation (RAG), ensuring quality through LLM-as-a-Judge filtering and conformal prediction calibrated by human expert annotations. An evaluation of five state-of-the-art LLMs on OmniVul reveals distinct performance gaps, with top-1 accuracy remaining below 50% on average for vulnerable code detection and CVE identification. Our evaluation also demonstrates that current models lack critical reasoning capabilities for reliable vulnerability assessment. These results highlight the importance of OmniVul for advancing research in evaluating and fine-tuning LLMs for vulnerability assessment.

Vishnu Teja Kandalam, Viet Duong, Xiaochang Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

From Threat Intelligence to Detection: Knowledge-driven Enrichment and Template-based Rule Grounding for Automated Sigma Rule Generation

AUTOSIGMA, an automated solution for transforming unstructured CTI reports into relevant Sigma rules that enables accurate, context-aware, and relevant rule generation, outperforms alternative solutions and LLM models in rule validity, rule relevancy, MITRE ATT&CK technique coverage, and robustness to input quality.

Sepehr Ghaffarzadegan, Boubakr Nour, M. Pourzandi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Staying on the Attack Path: Structured State for Long-Horizon Automated Penetration Testing

Large language model (LLM) based agents are increasingly applied to cybersecurity tasks such as vulnerability discovery and automated penetration testing. On long-horizon security tasks, however, such agents remain limited by context forgetting and intent drift: early critical facts and causal reasoning chains are lost over extended interactions, and the agent falls into aimless, repetitive exploration. This paper proposes Intentest, an intent-graph-guided automated penetration testing agent that externalizes long-horizon state from the LLM's context window onto a persistent fact-intent directed acyclic graph (DAG), thereby substantially reducing invalid transitions. We evaluate Intentest on automated penetration testing of web applications, a representative long-tail task in cybersecurity. In the DAG, verified network states are stored as immutable fact nodes, and exploration directions are constrained as intent edges bounded by predecessor facts. The system adopts a three-layer architecture, in which the fact-intent mapping layer maintains the global state, the task scheduling and allocation layer ensures execution stability through two-phase degradation recovery and multi-dimensional adaptive load balancing, and the intent retrieval and prediction layer provides tactical priors through a top-down five-stage filtering algorithm. On a benchmark of real CTF challenges covering more than ten vulnerability types across three difficulty levels, Intentest achieves an overall success rate of 88.2% and a success rate of 75.0% on hard tasks, improving over the baseline by approximately 44 and 50 percentage points. Ablation experiments further show that the intent retrieval and prediction reduce the average number of rounds on successful medium and hard tasks by about 33% and 48%, respectively, without changing the set of solvable tasks.

Wei-Zhe Wang, Yi-Tong Zhang, Yao Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.