This work introduces context segmentation, a two-level agentic framework that divides complex exploitation tasks into manageable, contextually isolated sub-problems and demonstrates that for the E4B model, the strategy acts as an intelligent search, achieving competitive rewards with superior token efficiency compared to brute-force retries.
Abstract
The proliferation of highly capable open-weight Small Language Models (SLMs) democratizes access to advanced cybersecurity capabilities, posing an escalating risk as these models can bypass proprietary API guardrails when deployed locally. However, SLMs deployed as autonomous agents often struggle with long-horizon, exploratory tasks like cybersecurity Capture The Flag (CTF) challenges due to context bloat and cognitive degradation from accumulated tool-call outputs. To understand and mitigate this cybersecurity threat, we introduce context segmentation, a two-level agentic framework that divides complex exploitation tasks into manageable, contextually isolated sub-problems. Evaluating on the picoCTF dataset using memory-constrained gemma-4 models, we demonstrate that for the E4B model, our strategy acts as an intelligent search, achieving competitive rewards with superior token efficiency compared to brute-force retries, and successfully solving 18.52% of tasks that standard agentic execution fails to complete. Code is available at https://github.com/9xeb/context-segmentation.
The findings show that LLMs can approximate structured cybersecurity reasoning under controlled representations, but do not apply it robustly, which has important implications for the design and evaluation of AI-assisted security decision-support systems.
Cybersecurity combines high-stakes analysis with complex technical language, making it an impactful and challenging domain for LLMs. We present MiST (Mid-trained Security Transformer), a suite of 8B and 32B models that achieve strong performance on public cybersecurity benchmarks. We use mid-training as an intermediate...
O. Ovadia, Elad Ben Zaken, Elad Guttman et al.· 0 citations
CyberFactory is introduced, a unified open-source framework that connects data construction, trajectory synthesis, and model training across proof-of-concept (PoC) generation, vulnerability patching, and cybersecurity question answering (CyberQA).
Jian Yang, Haau-Sing Li, Shawn Guo et al.· 0 citations
PrivEscalate is presented, a large-scale benchmark for Linux privilege escalation, comprising 531 Dockerized scenarios spanning 14 sub-categories and PrivEscalate, a domain-specialized wrapper that augments a generic ReAct agent with deterministic enumeration, category matching, and step planning that improves over pri...
Yi-Xuan Liu, Zi-Long Zhen, Yin Wu et al.· 0 citations
Translating high-level controls from security standards into concrete, system-specific requirements is central to cybersecurity requirements engineering. Large language models (LLMs) can accelerate this labor-intensive, recall-sensitive task, but any single run is unreliable: it misses valid safeguards while introducin...
Santiago Perez-Acuna, Y. Martín, J. Yelmo· 0 citations
This work presents an end-to-end diagnostic study of an Autonomous Adversary system with orchestrator, executor, and validator LLMs in enterprise-like lateral-movement scenarios and uses comparative LLM-as-a-Judge analysis to identify planning deficiencies.
Saeedeh Lohrasbi, Mohammad Mamun, Ahmed Yehia et al.· 0 citations