Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Sep 2026

Evidence-Grounded Retrieval for Investigation Hunt Lead Generation from CTI Reports

Threat hunting increasingly depends on converting unstructured knowledge (e.g., Cyber Threat Intelligence reports) into actionable hunt leads: concise, investigable hypotheses grounded in observable artifacts and adversary techniques. Producing such leads manually is a tedious and hard-to-scale task. Existing automated...

A. Prakash, Boubakr Nour, M. Pourzandi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks

Large language model (LLM) benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cybersecurity benchmarks across 10 proprietary, open-weight, and cybersecurity-specialized LLMs. By modeling benchmarks as measurement pipelines, w...

Aymene Berriche, Cathrine Shalby, Mohannad J. Alhanahnah et al. · 1 citation
#artificial intelligence Preprint Sep 2026

ACEA: An Adversarial Co-Evolution Arena for Head-to-Head Red-Team and Blue-Team LLM Testing

A platform that connects a pluggable red-team adapter and a pluggable blue-team adapter to a shared target LLM and scores their attack and defense rates with an LLM judge, and describes the design of ACEA and the metrics through which red and blue teams are scored head to head.

Yi-Da Shen, Kentaroh Toyoda, Alex Leung · 0 citations
#artificial intelligence Preprint Sep 2026

Do AI Coding Assistants Check Before They Install? A Pre-Registered Demand-Side Audit of Trust Signals in the Research Software Supply Chain

A controlled study on six open-source research software projects, with protocol, seed, panel, and analysis plan deposited with a DOI before any trial, drew three conclusions: publishing signals is necessary but not sufficient; price did not buy verification; verification must be built into the program that runs the ass...

Pengyin Shan · 0 citations
#artificial intelligence Preprint Sep 2026

Your Agent Says Yes: Interpreting Adversarial Market Behavior Beyond Individual Transactions

Repeated runs show that category-level and within-trajectory relations can recur even when normalized score-change rankings do not, and motivate agent-behavior evaluation that links communication, authorization, and evolving state instead of treating individual transaction verdicts as complete safety judgments.

Ze-Lin Li, Yi-Yun Su, Matt White et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Staying on the Attack Path: Structured State for Long-Horizon Automated Penetration Testing

Intentest is proposed, an intent-graph-guided automated penetration testing agent that externalizes long-horizon state from the LLM's context window onto a persistent fact-intent directed acyclic graph (DAG), thereby substantially reducing invalid transitions.

Wei-Zhe Wang, Yi-Tong Zhang, Yao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Towards a Resilience-Theoretic Foundation for Adversarial Robustness in Industrial Control System Anomaly Detection

It is established that adversarial robustness in ICS anomaly detection is a specific instantiation of system resilience, and a compositional resilience bound for heterogeneous ICS detection networks is derived, showing that the binding constraint on system-level resilience is the coupling-adjusted absorption capacity o...

Branka Stojanović, Andreas Flatscher, Michael Somma · 0 citations
#artificial intelligence Preprint Sep 2026

AgentLeak: Cloning Stronger LLM Agent Capabilities onto Weaker Agents Beyond Skill Stealing

The findings reveal a confidentiality risk in LLM agents: protecting explicit artifacts alone is insufficient, as observable execution behavior can leak the procedural knowledge required to reconstruct proprietary task-solving capabilities in low-capability and attacker-controlled agents.

Xiaoting Lyu, Yu-Hong Wu, Yu-Fei Han et al. · 1 citation
#artificial intelligence Preprint Sep 2026

CIPHER: Benchmarking Cross-record Inference over Privacy-Hardened Evidence Records

This work introduces CIPHER (Cross-record Inference over Privacy-Hardened Evidence Records), a benchmark of expert-validated questions from consumer-finance, clinical, and law-enforcement records that evaluate retrieval, prompting, table-specialist, and hybrid symbolic-neural systems under native redaction and surrogat...

Suparno Roy Chowdhury, M. Choudhury, Dhruv Madhwal et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AgentDrift: A Step-Labeled Benchmark of Injection-Hijacked LLM Agent Trajectories

AgentDrift is presented, a benchmark of 12,536 synthetic tool-call trajectories over five agent domains in which every one of the 71,024 steps carries one of four labels: benign, injection point, hijacked, or failed injection; it is shown that the LLM judge was itself fooled by the hard negatives.

Asif Pinjari, Mithun Paul Saint-Germain · 2 citations
#artificial intelligence Preprint Sep 2026

WAPP: Safe Learning of Positive Security WAF Policies from Live Traffic

Web Application Firewalls (WAFs) mainly rely on signatures to detect known attacks, which can leave gaps against modified or previously unseen payloads. Positive security provides a complementary approach by learning legitimate traffic and blocking inputs that fall outside the learned profile. However, learning directl...

H. Osama, Zeyad Ahmed, M. Amgad et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.