Skip to content

Evaluating Context Segmentation in Locally Deployable SLMs for Cybersecurity CTF Tasks

Sep 2026 · 0 citations · 25 references
Computer Science

TL;DR

This work introduces context segmentation, a two-level agentic framework that divides complex exploitation tasks into manageable, contextually isolated sub-problems and demonstrates that for the E4B model, the strategy acts as an intelligent search, achieving competitive rewards with superior token efficiency compared to brute-force retries.

Abstract

The proliferation of highly capable open-weight Small Language Models (SLMs) democratizes access to advanced cybersecurity capabilities, posing an escalating risk as these models can bypass proprietary API guardrails when deployed locally. However, SLMs deployed as autonomous agents often struggle with long-horizon, exploratory tasks like cybersecurity Capture The Flag (CTF) challenges due to context bloat and cognitive degradation from accumulated tool-call outputs. To understand and mitigate this cybersecurity threat, we introduce context segmentation, a two-level agentic framework that divides complex exploitation tasks into manageable, contextually isolated sub-problems. Evaluating on the picoCTF dataset using memory-constrained gemma-4 models, we demonstrate that for the E4B model, our strategy acts as an intelligent search, achieving competitive rewards with superior token efficiency compared to brute-force retries, and successfully solving 18.52% of tasks that standard agentic execution fails to complete. Code is available at https://github.com/9xeb/context-segmentation.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

MiST: Mid-Training LLMs for Cybersecurity

Cybersecurity combines high-stakes analysis with complex technical language, making it an impactful and challenging domain for LLMs. We present MiST (Mid-trained Security Transformer), a suite of 8B and 32B models that achieve strong performance on public cybersecurity benchmarks. We use mid-training as an intermediate...

O. Ovadia, Elad Ben Zaken, Elad Guttman et al. · 0 citations
#small language model Preprint Sep 2026

PrivEscalate: Measuring and Augmenting the Threat of LLM-Automated Linux Privilege Escalation

PrivEscalate is presented, a large-scale benchmark for Linux privilege escalation, comprising 531 Dockerized scenarios spanning 14 sub-categories and PrivEscalate, a domain-specialized wrapper that augments a generic ReAct agent with deterministic enumeration, category matching, and step planning that improves over pri...

Yi-Xuan Liu, Zi-Long Zhen, Yin Wu et al. · 0 citations
Review Sep 2026

Ensembling LLMs for AI-Augmented Cybersecurity Software Requirements Generation

Translating high-level controls from security standards into concrete, system-specific requirements is central to cybersecurity requirements engineering. Large language models (LLMs) can accelerate this labor-intensive, recall-sensitive task, but any single run is unreliable: it misses valid safeguards while introducin...

Santiago Perez-Acuna, Y. Martín, J. Yelmo · 0 citations
#artificial intelligence Preprint Sep 2026

Where Cyber Agents Struggle: Bottleneck Analysis of Multi-Stage LLM Agents

This work presents an end-to-end diagnostic study of an Autonomous Adversary system with orchestrator, executor, and validator LLMs in enterprise-like lateral-movement scenarios and uses comparative LLM-as-a-Judge analysis to identify planning deficiencies.

Saeedeh Lohrasbi, Mohammad Mamun, Ahmed Yehia et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.