Skip to content

Automating Attack Graph Construction for Agentic Pentesting. Towards Neuro-Symbolic Vulnerability Hunting

Sep 2026 · 0 citations · 24 references
Computer Science

TL;DR

A semi-automated pipeline is presented that addresses the interoperability problem of integrating symbolic frameworks such as MulVAL to contemporary security workflows or agentic pipelines, but predicate coverage, rule coverage, and path precision remain limiting factors.

Abstract

Logic attack graphs grounded in scanner output provide explicit and auditable attack path reasoning LLM-based agents lack. Integrating symbolic frameworks such as MulVAL to contemporary security workflows or agentic pipelines, however, requires translating scanner evidence to initial facts, and creating domain-specific rules. We present a semi-automated pipeline that addresses this interoperability problem and depict its feasibility in a web-security case study. Our pipeline parses findings from Trivy, Semgrep, and Nmap into MulVAL predicates and uses an LLM-assisted process to construct domain-specific Datalog rules linking scanner-detectable evidence to attack techniques. MulVAL/XSB then performs symbolic inference to generate structured attack paths. We evaluate the attack-graph construction infrastructure on 54 web Capture-the-Flag tasks from CyBench within an agentic pipeline (Hybrid Reasoner); we do not evaluate the performance of the downstream agent. Every task produced at least one goal-reaching graph, and we achieve mean ground-truth vulnerability coverage of 53.7%, with 51.9% achieving full coverage; mean noise-path rate was 83.9%. With median end-to-end time of 24.9 s (MulVAL reasoning: 2.7 s) the pipeline is feasible and runtime-practical for agentic workflows, but predicate coverage, rule coverage, and path precision remain limiting factors. Next steps include semantic rule validation and agent-level comparison for graph-guided pentesting.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance

Large language model agents have demonstrated promising capabilities in cybersecurity tasks, yet their ability to reconstruct complete Advanced Persistent Threat attack campaigns from complex security logs remains largely unexplored. Existing cybersecurity benchmarks for agents mainly focus on vulnerability discovery,...

Qi Chen, Fu-Shuo Huo, Hang-Li Shen et al. · 0 citations
#artificial intelligence Preprint Oct 2026

CVE2AP: Automated Generation of PDDL-Encoded Attack Paths via Large Language Models

Attack Path (AP) modeling is fundamental to cybersecurity analysis, where the Planning Domain Definition Language (PDDL) has been widely adopted to encode APs into formal and machine-verifiable representations for automated reasoning about vulnerability exploitation, attack progression, and their potential impacts. How...

Lin Cui, Vincenzo Scotti, R. Mirandola · 0 citations
Review Open access 2026

A Multi-Agent DevSecOps Framework for Intelligent Vulnerability Detection and Auto-Remediation

This paper presents a multi-agent DevSecOps framework that integrates static code scanning, large language model (LLM) based security reasoning, automated repair generation, policy-as-code enforcement, and runtime monitoring into a unified event-driven pipeline. Five specialized agents collaborate through LangGraph sha...

Hai-Ning Fan, Li-Jie Zheng, Chen-Hao Han et al. · 0 citations
Sep 2026

SE4SC-LLM: an LLM-Augmented symbolic execution framework for smart contracts

SE4SC-LLM, an LLM-augmented symbolic execution framework for smart contracts that achieves 95.1% average CFG coverage, a 6.5 percentage point improvement over the strongest baseline, and detects 11.2% more vulnerabilities.

Tian-Huan Miao, Yang Liu · 0 citations
Preprint Aug 2026

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

RedAgentBench is introduced, an executable framework for autonomous red-teaming and faithful measurement that shows that executable evaluation can improve safety measurement and identify actionable intervention points.

Zixing Chen, Xingyuan Liu, Jie Zhu et al. · 3 citations
#artificial intelligence Preprint Sep 2026

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

This work presents AgentXploit, a two-role auditing system that separates repository-level attack-path discovery from runtime exploitation and introduces AgentXploit-Bench, containing 72 reproducible vulnerabilities across 12 open-source AI-agent systems and frameworks.

Wei-Da Liang, Shi Qiu, Zhun Wang et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.