Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Open access Sep 2026

Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark

With LLM watermarking already being deployed commercially, practical applications increasingly require multibit watermarks that encode more complex payloads, such as user IDs or timestamps, into the generated text. In this work, we propose a fundamentally new approach for multibit watermarking: introducing binomial enc...

Thibaud Gloaguen, Robin Staab, Mark Vero et al. · 0 citations
#artificial intelligence Review May 2026

FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence

FraudBench is a multimodal benchmark for detecting AI-generated fraudulent refund evidence and shows that current MLLMs often recognize real-damaged evidence but fail on many fake-damaged subsets, with fake-damage detection rates far below the 50\% baseline on most generator subsets.

Xinyu Yan, Bo-Yang Chen, Jia-Ming Zhang et al. · 1 citation
#artificial intelligence Review May 2026

MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents

MOSAIC-Bench (Malicious Objectives Sequenced As Innocuous Compliance), a benchmark of 199 three-stage attack chains paired with deterministic exploit oracles on deployed software substrates that treats both exploit ground truth and downstream reviewer protocol as first-class evaluation axes, is introduced.

J. Steinberg, Oren Gal · 4 citations
#artificial intelligence Preprint Nov 2025

Server-Enforced Watermarking in U-Shaped Split Federated Learning

Sigil, a server-enforced watermarking framework for U-SFL, defines a secret watermark constraint in the server-visible activation space and embeds the watermark into client-side models by injecting a watermark gradient into the gradients returned during training.

Zhengchunmin Dai, Jia-Xiong Tang, Peng Sun et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Assuming You Knew: Fixing an Epistemic Semantics for Flow Policies Using Agentic AI

Many high-level security requirements are about the allowed flow of information in programs and are difficult to make precise because they involve selective downgrading. Notions from epistemic logic have emerged as a good approach to policy semantics but a robust general framework remains elusive. A paper appearing in...

David A. Naumann · 0 citations

Enabling Regulatory Multi-Agent Collaboration: Architecture, Challenges, and Solutions

This paper proposes a blockchain-enabled layered architecture for regulatory agent collaboration, comprising an agent layer, an off-chain computation layer, and an on-chain anchoring layer that establishes a systematic foundation for trustworthy, resilient, and scalable regulatory mechanisms in large-scale agent ecosys...

Qin-Nan Hu, Yuntao Wang, Yuan Gao et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Distillation Defenses Easily Break After Reinforcement Learning

Distillation attacks copy the reasoning capabilities of closed-source large language models, allowing bad actors to replicate state-of-the-art performance at low cost. Attackers systematically collect a large volume of frontier model reasoning traces and then train (i.e.,"distill") their own models on these traces. Exi...

Shidan Javaheri, Alexander Panfilov, O. Britton et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents

This work introduces SEABench, a benchmark for studying endogenous misalignment arising from agent self-evolution, with 48 longitudinal task sequences that span multiple evolution surfaces, task domains, and harm types in a rich personal-assistant environment and shows that qualitatively different safety behaviors emer...

Saswat Das, Parvati Viswanathan, Daniel Donnelly et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Compiler May Read It, the Agent May Not: Keeping Part of a Research Code Away from a Coding Agent

The compiler must read modules a physics-based solver cannot build without; the coding agent must not read that intellectual property. The harness does not ship that rule. We classified fifteen read routes against a container, permission rules and a sandbox. None of the three can tell which program is reading.

Shobhan Roy · 0 citations
#artificial intelligence Preprint Open access Sep 2026

"Nothing to See Here'': Unintended Disclosure through Revision Traces of LLM Deliverables

Large language model (LLM) assistants increasingly help users draft content for third-party recipients. During private drafting, the user or the model may introduce an item and later remove or replace it. The model may remove the item from the intended content but reveal it again when stating the edit. We call such sta...

Yage Zhang, Yukun Jiang, Yang Zhang · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Planarian: Managing Agent State with Statepoints

LLM agents solve complex tasks by iteratively changing files, invoking local tools, and interacting with remote services, which modifies state across their local environment and remote services. Today, agents and users must manage these changes explicitly, whether reverting exploratory actions or recovering from errone...

Jinnan Guo, Hao Mark Chen, Kapil Vaswani et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.