Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Preprint Open access Oct 2026

Can Agents Trust Their Skills? Uncovering Unsafe Chains of Trust in Skill-Based LLM Agents

LLM agents increasingly rely on installable skills, which are packages of instructions, code, and resources that equip them with task-specific capabilities and, once installed, can be automatically invoked across subsequent user tasks. This creates a chain of trust in which users delegate authority to agents, while age...

Yan Wang, Zhihao Zhang, Ke Chen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Approval Laundering: Systematizing Approval--Execution Binding Failures in AI Coding-Agent Harnesses

Modern AI coding-agent harnesses (Claude Code, Codex CLI, Cursor) rest their security boundary on a largely unexamined assumption: that the action A a human approves is the same action A'the harness executes, where A is fixed by a stated policy for what a scope grant or session-scoped approval authorizes. We show this...

Yang Wang · 0 citations
#artificial intelligence Preprint Sep 2026

APTInvestBench: Evaluating Autonomous APT Investigation under Varying Telemetry

Large language model (LLM) agents could help security operations centers (SOCs) investigate advanced persistent threats (APTs) by turning weak leads into evidence for intrusion scoping and response. Yet success under one telemetry setting does not establish robustness to changes in log collection, retention, or samplin...

Yu Wang, Shu-Hao Li, Tao Yin et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SceneJail: Exploiting Video Scenario Context to Jailbreak Multimodal LLMs

Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby treating video merely as a carrier. Consequently...

Wenyu Chen, Li Wang, Chuanchao Zang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Alignment via Training Against Probes Without Losing Monitorability

Models are usually aligned based on their observed outputs, using demonstrations, preference data, or reward signals. These objectives reward responses that look aligned. More capable models may learn to satisfy them without internalizing the intended behavior, for example by faking compliance during training. Such sup...

Lena Libon, Alexander Panfilov, Ben Rank et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Security-Enhanced Seed-Based Weight Quantization for Large Language Models

Large language models (LLMs) incur substantial storage, memory-bandwidth and energy costs, motivating compact weight representations. Existing seed-based compression methods reconstruct weights from compact pseudo-random representations but do not explicitly account for the non-uniform sensitivity of model weights. We...

Qiu-Yu Ren, Sudipta Paria, Aritra Dasgupta et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks

This technical report presents an alignment evaluation developed and performed by the UK AI Security Institute for assessing whether advanced AI systems take unsanctioned actions outside the scope of their assigned task. We evaluate whether frontier models conduct supply-chain attacks against out-of-scope, third-party...

Alexandra Souly, Kai Fronsdal, Abby D'Cruz et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LogiC-Diff: Embedding Security Properties Into AI-Enabled Cyber-Physical Systems

AI-enabled Cyber-Physical Systems (CPS) are highly vulnerable to adversarial and anomalous inputs, where small perturbations can induce cascading errors and unsafe control actions. Existing approaches, such as rule-based filtering, training-time regularization, or diffusion-based reconstruction, either operate outside...

Zi-Yan An, John Stankovic, Mei-Yi Ma · 0 citations
#artificial intelligence Preprint Sep 2026

Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning

Federated learning (FL) has become a foundational paradigm for multi-institutional medical AI, allowing hospitals and research centers to jointly train diagnostic models without exchanging patient records. This privacy promise, however, is increasingly contested: a malicious or honest-but-curious server can launch mode...

Chao-Yu Zhang, Shang-Hao Shi, Heng Jin et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Janus: Evidence-Before-Effect Sagas and Offline-Verifiable Provenance for Agentic LLMs

Agentic large language models (LLMs) now move money through tools, yet the record of what they did is usually a trace their own process emits beside the effect. Janus puts the record on the effect path. A step's proposal, the verdict on it and any answer from a validator or a person are durable in a signed, hash-chaine...

Mustafa Arslan · 0 citations
#artificial intelligence Preprint Sep 2026

CollageAttack: Exploiting Cross-Modal Alignment Flaws in T2I Models through Spatial Text Composition

Text-to-image (T2I) models have substantially improved in language understanding, in-image text rendering, and visual composition, while their safety mechanisms do not always keep pace with these capabilities. This creates a cross-modal attack surface in which harmful semantics can remain inconspicuous in a serialized...

Zhi-Yi Mou, Yao Lu, Wang-Ze Ni et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Forensic-Aware Continual Adaptation for Image Forgery Localization

The rapid evolution of image manipulation techniques has raised growing public security concerns. Existing Image Forgery Localization (IFL) methods can accurately localize manipulated regions but are often unable to adapt to newly emerging forgeries. In real-world forensic scenarios, data typically arrive sequentially,...

Chen-Qi Kong, Song Xia, An-Wei Luo et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.