2026· International Journal of Advanced Computer Science and Applications· Vol 17· 0 citations· 14 references
TL;DR
This study evalu-ates zero-shot and few-shot prompting conditions and studies the effect of iterative prompt refinement, focusing on explicit format constraints for cryptographic hash entities.
Abstract
Cyber Threat Intelligence reports combine analytical prose with dense technical indicators, making structured entity extraction a challenging but operationally valuable task. This study presents a comparative evaluation of three large language models – Claude Sonnet 4.6, GPT-5.4, and LLaMA 4 Scout – on a manu-ally annotated corpus of 21 real-world CTI reports across 15 entity types and 1284 ground truth instances. This study evalu-ates zero-shot and few-shot prompting conditions and studies the effect of iterative prompt refinement, focusing on explicit format constraints for cryptographic hash entities. Results show that Claude Sonnet 4.6 and GPT-5.4 achieve comparable perfor-mance under zero-shot conditions, with LLaMA 4 Scout trailing by a substantial margin. Few-shot prompting consistently reduc-es hallucination rates, but yields mixed F1 results, with exemplar cardinality emerging as a critical and underappreciated design factor. Entity extraction difficulty varies substantially across types, with technical indicator categories showing near-perfect performance and semantic categories such as tool and target sector posing the greatest challenges across all evaluated mod-els.
STINER, a taxonomy and expert-annotated corpus for extracting strategic intelligence from social media streams is introduced, and how social-media-driven extraction can surface early signals of the SafePay ransomware campaign prior to its retrospective characterization in vendor threat landscape reports is illustrated.
Yasir Ech-Chammakhy, Oussama Azrara, J. Chbili et al.· 0 citations
Automated extraction of structured threat information from unstructured cyber threat intelligence (CTI) underpins modern security operations, yet the supporting machine learning resources are almost exclusively English: no annotated Arabic CTI named entity recognition (NER) corpus has been published. We introduce AraCTI-NER, a dataset of 10,312 token-level annotated samples (275,530 tokens; 42,360 entity spans) over eight STIX-inspired entity types, built by an LLM-assisted pipeline seeded with authentic Arabic cybersecurity articles, structurally validated and rebalanced through targeted generation. We benchmark seven encoders from three families (Arabic-specialized, English cybersecurity-adapted, and multilingual) over three seeds under strict entity-level metrics, and release a 408-sentence expert-audited test subset (ATS-gold) whose reliability is quantified by a second independent expert validation (inter-annotator agreement 0.878 entity-level F1). XLM-RoBERTa Large attains the best mean F1 (0.7603; 0.7674 on ATS-gold), with AraBERTv2 close behind (0.7491), while both English-only cybersecurity encoders fall to ≈0.63, a separation that holds across every seed and survives expert correction, with the ≈3-point F1 decrease from ATS-silver to ATS-gold concentrated in Vulnerability and TTP. On 350 doubly annotated sentences from authentic Arabic cyber-incident news, a shift in both provenance and register, the strongest model reaches F1 = 0.5429 against an inter-annotator F1 of 0.616. AraCTI-NER establishes the first reproducible baseline for Arabic CTI NER and identifies domain-adaptive Arabic cybersecurity pre-training as the highest-value next step.
CyberNER is introduced, a two-stage pipeline to solve the multi-type NER and alias canonicalization problem in APT CTI reports and achieves a Macro-F1 score of 0.853, outperforming all four baselines.
Unnamalai K, Suriakala M· International journal of com...· 0 citations
A manually constructed dataset of 150 English-language CTI reports, each represented as STIX 2.1 based graphs, provides a benchmark for CTI information extraction, knowledge-graph construction, incident analysis, and threat attribution and indicates that locally deployed LLMs can support human reviewers in identifying annotation inconsistencies, but expert validation remains essential.
Dipshikha Das, Arnab Banik, Md. Shariful Islam et al.· arXiv.org· 0 citations
The threat posed by adversarial prompts to large language models is becoming harder to ignore. Problems including prompt injection, jailbreaking, phishing, and Unicode-based attacks are now widespread. Most existing solutions protect against only one threat type, operate in English only, and provide no explanation for their decisions. We present SemGuard, a multilingual security gateway using Triple-Anchor Semantic Threat Modeling, which simultaneously evaluates each input against three semantic reference sets: attack, safe, and destructive. SemGuard detects four threat types concurrently in Arabic, Arabizi, and English. We expand the original Arabic Security Dataset from 319 to 807 validated examples across seven threat categories, using three independent LLM judges (GPT-4o, Grok-4, Llama 3.3 70B) achieving Fleiss' $\kappa=0.839$. After retraining on the expanded dataset, SemGuard achieves a mean F1-score of 0.989 and recall of 0.991, representing a 13.7% improvement over the original implementation. Analysis of 527 rejected examples reveals quantitative evidence of threat-category ambiguity, with impersonation exhibiting a 98.2% inter-judge disagreement rate, validating the necessity of the Triple-Anchor framework. This work also presents the first Arabic LLM security dataset with a formal LLM-as-Judge annotation protocol.
Abdullah M. Abughallous, Somia Abufakher· IEEE Jordan Conference on Ap...· 0 citations
Automated parsing of Cyber Threat Intelligence (CTI) is crucial for threat attribution and proactive defense. However, Chinese CTI texts are highly unstructured and semantically fragmented, posing dual challenges for existing models. In entity extraction, fragmented tokenization caused by high-entropy entities and complex nested structures leads to ambiguous entity boundaries. In relation extraction, critical attack clues are scattered across paragraphs, preventing traditional attention mechanisms from effectively capturing long-range dependencies and relative spatial structures. To address these limitations, we propose an adaptive global-local joint extraction framework designed for fragmented semantic aggregation in Chinese CTI. Within the entity recognition module, we introduce adaptive rotary position embeddings to correct low-level positional features. This mechanism, combined with a type-decoupled GlobalPointer, resolves recognition conflicts involving long-span entities and nested boundaries. In the relation extraction module, we design a dual-stage attention mechanism to dynamically integrate global cross-paragraph spatial clues with local entity neighborhood features. Additionally, an adaptive decoding strategy aware of class imbalance is implemented to enhance the robustness of the model against sparse long-tail relations. Experimental results on the CDTier dataset indicate that the proposed framework achieves a 9.45% improvement in the F1 score of entity extraction over the best existing baseline, alongside a precision of 93.3% and a recall of 95.5% for relation extraction. The proposed method overcomes the bottleneck of long-range semantic parsing in complex Chinese contexts, demonstrating superior generalization capabilities and practical utility.
Jipeng Tang· Poster Volume 0008 The 2026...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.