Skip to content

GARAGE: Characterizing the Automation Boundary in LLM-based Attack Graph Generation

Jul 2026 · arXiv.org · Vol abs/2607.18108 · 0 citations · 28 references
Computer Science

TL;DR

GARAGE is introduced, a RAG-powered framework that converts fragmented CTI into an actionable, domain-specific knowledge base for automated attack graph generation and position GARAGE as a scalable TARA support tool within human-in-the-loop workflows, offering a comprehensive cost-performance analysis to guide its deployment across various LLM tiers.

Abstract

While modern vehicle security depends on effective Cyber Threat Intelligence (CTI) synthesis, current automated tools struggle with unstructured data and automotive-specific architectural nuances. To bridge this gap, we introduce GARAGE, a RAG-powered framework that converts fragmented CTI into an actionable, domain-specific knowledge base for automated attack graph generation. GARAGE synthesizes a dataset of 12,786 CVEs and 140 incident reports into a STIX 2.1 and Auto-ISAC ATM-compliant knowledge base. By formalizing tactical-pattern-level scenarios through granular kill chain analysis, GARAGE achieves threat generation capabilities. Our 320 Leave-One-Out experiments reveal that the framework can accurately transfer security knowledge to entirely unseen vehicle architectures. Furthermore, we position GARAGE as a scalable TARA support tool within human-in-the-loop workflows, offering a comprehensive cost-performance analysis to guide its deployment across various LLM tiers.

View source

Similar papers

Review Open access Aug 2026

From static tasks to dynamic reasoning: a characterization framework and study of large language models in next-generation cybersecurity automation

This work surveys recent LLM-based systems across seven core domains and identifies the need for privacy-aware deployment, timely retrieval and knowledge maintenance for emerging threats, process-level evaluation tied to measurable security outcomes, and human oversight within controlled and hybrid automation workflows.

Hanxin Yu, Shahrear Iqbal, E. C. Pinto et al. · 0 citations
Conference Open access 2026

Graph2TTP: Knowledge Graph-Guided Paragraph-Level TTPs Identification from Cyber Threat Intelligence Reports

Graph2TTP is proposed, a novel neural-symbolic framework for automated, paragraph-level Tactic, Technique and Procedure (TTP) identification that outperforms state-of-the-art neural baselines and establishes a robust new standard for accurate and interpretable threat intelligence analysis.

Patrick Zounon, Yu-Fei Han, Michel Hurfin et al. · 0 citations
Review Open access Aug 2026

CAPS: Compositional Attack Path Scoring for LLM Deployment Stacks

Compositional Attack Path Scoring (CAPS), a framework engineered to quantify end-to-end multi-hop risks in LLM architectures, establishes a rigorous benchmark for quantitative vulnerability management in complex, agentic LLM environments.

Quang-Vinh Dang, Hoang-Viet Vu, Ngoc-Son-An Nguyen et al. · 0 citations
Preprint Jul 2026

Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of Predicate Representation Granularity

This work investigates whether AURORA's nine-category taxonomy provides representational distinctions beyond those captured by a reduced, empirically derived scheme, and suggests that higher granularity primarily enhances the internal structural resolution of a plan's justification rather than the viability of the generated attack chain itself.

Ramya Varunsegar · 0 citations
#natural language process... Preprint Aug 2026

BEACON: Behavior-Anchored Cross-Source Knowledge Graph Construction for Cyber Threat Intelligence

BEACON is an LLM-driven framework for cross-source CTI knowledge graph construction that constructs and releases two human-annotated datasets from 34 sources and outperforms all baselines by at least 23% and 9%, respectively.

Changze Li, Yutong Cheng, Tsania Camila Finnisa et al. · 0 citations
Book Open access Aug 2026

OmniVul: A Holistic, Multi-Turn Conversational Benchmark for LLM-Based Vulnerability Assessment

With more than 20,000 Common Vulnerabilities and Exposures (CVEs) reported annually, software vulnerabilities represent a critical cybersecurity challenge. This volume has intensified the demand for automated detection and analysis, motivating the integration of large language models (LLMs) for such tasks. However, existing vulnerability benchmarks are not suitable for evaluating LLMs' capabilities in vulnerability assessment, as most of them 1) rely on narrow data sources, 2) lack deep context, and 3) focus on single-turn Q&A rather than realistic, multi-stage analyst workflows. To address this gap, we introduce OmniVul, a comprehensive multi-turn benchmark for LLM-based vulnerability assessment. OmniVul comprises 2,000 CVEs with question–answer pairs spanning 23 attributes, including detection, code localization, root cause analysis, and patch suggestion. We employ an automated workflow to aggregate multi-source data via Retrieval-Augmented Generation (RAG), ensuring quality through LLM-as-a-Judge filtering and conformal prediction calibrated by human expert annotations. An evaluation of five state-of-the-art LLMs on OmniVul reveals distinct performance gaps, with top-1 accuracy remaining below 50% on average for vulnerable code detection and CVE identification. Our evaluation also demonstrates that current models lack critical reasoning capabilities for reliable vulnerability assessment. These results highlight the importance of OmniVul for advancing research in evaluating and fine-tuning LLMs for vulnerability assessment.

Vishnu Teja Kandalam, Viet Duong, Xiaochang Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.