Skip to content
Conference

Improving TTP Mapping Accuracy in CTI Reports Using Structured RAG Query Configuration

Jul 2026 · International Conference on Ubiquitous and Future Networks · pp. 465-470 · 0 citations · 19 references

Abstract

Accurately mapping attack behaviors described in Cyber Threat Intelligence (CTI) reports to the Tactics, Techniques, and Procedures (TTPs) of the MITRE ATT&CK framework is a critical challenge for responding to cyber threats and enhancing cyber resilience. However, traditional Large Language Model (LLM) and Retrieval-Augmented Generation (RAG) approaches face significant limitations. Specifically, the simple segmentation of unstructured CTI text leads to context fragmentation and the inclusion of extraneous descriptive details, which ultimately degrades the accuracy of TTP mapping. To solve these limitations, this paper proposes a methodology for constructing RAG queries based on structured fields. We extract attack behaviors from CTI reports as discrete events and organize them into seven fields (four required and three optional) representing the core components of the attack, which are designed to facilitate effective mapping to the MITRE ATT&CK framework. Our approach utilizes an LLM to perform event extraction and constructs optimized RAG queries through field combinations, thereby enhancing semantic alignment during the retrieval process. Experimental results demonstrate that the proposed method improves the F1-score by 0.249 compared to the baseline TTPFShot, achieving a maximum F1-score of 0.489 with the domain-specific model SecureBERT2. Moreover, while structured queries improve precision by constraining the retrieval scope to mitigate retrieval noise and reduce false positives, the effect of additional fields varies depending on the characteristics of the embedding model. Specifically, domain-specific models reach their peak performance with concise field configurations, whereas general-purpose models peak with the configuration integrating all fields. These results indicate that designing field combinations suited to the characteristics of the embedding model is essential, and that the proposed methodology provides a practical framework for high-precision TTP identification in complex CTI environments.

View source

Similar papers

#natural language process... Preprint Aug 2026

BEACON: Behavior-Anchored Cross-Source Knowledge Graph Construction for Cyber Threat Intelligence

BEACON is an LLM-driven framework for cross-source CTI knowledge graph construction that constructs and releases two human-annotated datasets from 34 sources and outperforms all baselines by at least 23% and 9%, respectively.

Changze Li, Yutong Cheng, Tsania Camila Finnisa et al. · 0 citations
Review Jul 2026

A Structured Cyber Threat Intelligence Dataset Using STIX 2.1 Entities and MITRE ATT&CK Mappings

A manually constructed dataset of 150 English-language CTI reports, each represented as STIX 2.1 based graphs, provides a benchmark for CTI information extraction, knowledge-graph construction, incident analysis, and threat attribution and indicates that locally deployed LLMs can support human reviewers in identifying annotation inconsistencies, but expert validation remains essential.

Dipshikha Das, Arnab Banik, Md. Shariful Islam et al. · 0 citations
Conference Open access 2026

Graph2TTP: Knowledge Graph-Guided Paragraph-Level TTPs Identification from Cyber Threat Intelligence Reports

Graph2TTP is proposed, a novel neural-symbolic framework for automated, paragraph-level Tactic, Technique and Procedure (TTP) identification that outperforms state-of-the-art neural baselines and establishes a robust new standard for accurate and interpretable threat intelligence analysis.

Patrick Zounon, Yu-Fei Han, Michel Hurfin et al. · 0 citations
#natural language process... Preprint Aug 2026

DisCTI: Who Needs to Know Timely? Automated Sector-Aware Cyber Threat Intelligence Dissemination

This work forms sector-targeted CTI dissemination as a multilabel classification problem, leveraging deep field knowledge of CTI structures and sector-specific threat patterns, and applies BERT, a transformer-based model, to automate the mapping of CTI events to sectors.

Fajar Wijitrisnanto, A. Abuadbba, Yan-Song Gao et al. · 0 citations
#small language model Review Sep 2026

A SoK for SoCs: Reading the TI Leaves on AI for Cyber Threat Intelligence Generation and Sharing

Three research directions for automating the production of shareable intelligence are derived from a literature survey of academic papers, organizing the CTI lifecycle into three stages: Threat Data Collection, CTI Generation and Sharing, and CTI Consumption.

Saastha Vasan, Hadjer Benkraouda, Jizhou Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.