Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Aug 2026

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity

A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving surface form varies, does the canonical-form score estimate model behavior well, and how much...

Yongxin Zhou, Jun-Wei Yao, Yuanzhe Liu et al. · 1 citation

An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic

This work forms model extraction monitoring as benign-calibrated traffic-window distribution testing: embed incoming queries into a semantic space and test whether their aggregate distribution deviates from historical benign traffic.

Shu-Ze Liu, Qian-Wen Guo, Yushun Dong · 1 citation

Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models

A novel function hijacking attack (FHA) that manipulates the tool selection process of agentic models to force the invocation of an attacker-chosen function and is largely agnostic to the context semantics and remains effective across domains and function sets.

Yannis Belkhiter, Giulio Zizzo, Sergio Maffeis et al. · 1 citation

T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search

A trajectory-aware evolutionary search method, T-MAP, which leverages execution trajectories to guide the discovery of adversarial prompts and enables the automatic generation of attacks that not only bypass safety guardrails but also reliably realize harmful objectives through actual tool interactions.

Hyomin Lee, Sangwoo Park, Yumin Choi et al. · 8 citations · ⚡1

More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles

This work proposes a general framework that utilizes weaker single-layer watermarks to preserve the entropy required for effective multi-layer ensembling and demonstrates that this counter-intuitive strategy mitigates signal decay and consistently outperforms strong baselines in both detectability and robustness.

Ruibo Chen, Yihan Wu, Xuehao Cui et al. · 2 citations

Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs

This work introduces a novel approach, \textit{CanaryTrace}, to safeguard the ownership of text datasets and effectively detect unauthorized use by RA-LLMs, and demonstrates high query efficiency, detectability, and consistency, along with minimal perturbation to the original dataset, all without compromising the perfo...

Yepeng Liu, Xuandong Zhao, D. Song et al. · 15 citations · ⚡1

Clinically Grounded Privacy Evaluation of Medical LMs

This work introduces a clinically grounded framework that evaluates leakage along a graded axis of adversarial access, ranging from publicly inferable demographics to leaked note fragments, and provides a practical, reusable framework for contextual privacy evaluation of medical LMs.

Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle et al. · 0 citations
#natural language process... Preprint Aug 2026

The Fragility of Jailbreak Robustness Across Operational States

This work finds that jailbreak robustness is highly fragile to operational-state variation: even when the attack remains fixed, changing only an ordinary system prompt not designed to affect safety can dramatically alter attack success rates.

Yuna Park, Hwang Youn Kim, Yujin Kim et al. · 0 citations
#natural language process... Preprint Aug 2026

Beyond the Payload: How User Invocation Shapes Coding Agent Vulnerability to Repository Poisoning

CIPR (Coding In Poisoned Repos), the first benchmark that systematically varies PLCs in poisoned real-world repositories, is introduced and highlights that coding agent vulnerability is not a static property, but a dynamic outcome shaped by everyday user configurations.

Fu-Kang Zhu, Bin-Bin Zhao, Rui-Xiao Lin et al. · 0 citations
#natural language process... Open access Aug 2026

OASIS: Optimizing Attacker Sequences for Hard-Label Black-Box Text Attacks

Experiments across multiple datasets, victim models, and large language models show that OASIS consistently outperforms strong standalone baselines and simple manually constructed chains, suggesting that attacker composition is not merely an implementation choice, but a practical optimization target for improving hard-...

Qian Chen, Shi-Liang Xiao, Yu-Zhi Liang · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.