Skip to content

Category

cybersecurity

1,065 papers

#artificial intelligence Review Sep 2026

Safety in Self-Evolving Agents: A Survey

Large language models (LLMs) exhibit strong general capabilities, yet their parameters typically remain fixed after deployment, limiting learning from new interactions. In open-ended environments, this motivates self-evolving agents that continually update reusable state-including model parameters, memories, tool defin...

Jia-Hao Chen, Zhou Feng, Ou-Bo Ma et al. · 1 citation
#artificial intelligence Preprint Oct 2026

Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, preserving data confidentiality of cloud computations. However, applying FHE to reinforcemen...

Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Can AI Oversight Be Zero Knowledge?

AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studi...

Alessandro Chiesa, Ziyi Guan, Burcu Yildiz · 0 citations
#artificial intelligence Preprint Sep 2026

Sapien: A Stateful Policy Engine for Autonomous AI Agents

Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing...

CO Tiffany, Wen Zhang, E. Bagdasarian et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection

With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean reference...

Jian-Wei Li, Jung-Eun Kim · 0 citations
#artificial intelligence Preprint Sep 2026

Backdoor Containment via Expert Quarantine and Shutdown in LLMs

Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning...

Jian-Wei Li, Min-Seon Kim, Jung-Eun Kim · 0 citations
#artificial intelligence Preprint Sep 2026

Fault-Tolerant Budget Conservation in Distributed Multi-Agent Delegation

An indistinguishability result shows that partition-local availability requires exclusive preallocation, and it is proved that ownership partition, ledger and effect conservation, descendant non-amplification, at-most-once settlement, late-completion safety, and partition confinement under explicit mediation, durabilit...

Gen-Liang Zhu, Chu Wang · 0 citations
#natural language process... Preprint Sep 2026

Anchor-ECC: Local Integrity Checking for Watermarked LLM Outputs via Error-Correcting Codes

LLM watermarking has become an effective approach to distinguishing AI-generated text from human-written text by embedding detectable patterns during generation. However, a small post-generation edit may change the meaning of the text without removing its overall watermark signal, creating a risk that the modified cont...

Ze-Wei Deng, M. Siddeek, Li-Yan Xie et al. · 0 citations
#natural language process... Preprint Open access Oct 2026

SecureVibe: Making Vibe Coding More Secure

As vibe coding becomes increasingly capable and widespread, security vulnerabilities in even functionally correct solutions are a growing concern. When investigating functionally correct but insecure solutions, we find that the insecure agent is less than half as likely to conduct effective planning and testing for the...

Danqing Wang, Baolin Peng, Zhepei Wei et al. · 0 citations
#natural language process... Preprint Sep 2026

The Geometry of Harmfulness in Multi-Turn Attacks

Large language models (LLMs) remain vulnerable to adversarial attacks that circumvent safety alignment to elicit harmful outputs. It remains unclear how harmfulness and refusal representations evolve over the course of multi-turn attacks, and why single-turn defenses are less effective in multi-turn settings. This work...

Yelyzaveta Husieva, Lauren Alvarez · 0 citations
#natural language process... Preprint Sep 2026

HARDE: Optimizing Agent Harnesses for Runtime Risk Detection and Execution Control

Large language model (LLM) agents are vulnerable to safety risks such as injected malicious instructions or misleading information, motivating runtime defenses that prevent unsafe action in execution across diverse risks while preserving benign-task utility. Existing system-level defenses either focus on risk detection...

Zhuo Liu, Mo-Xin Li, Zhi-Xin Ma et al. · 0 citations
#cybersecurity Book Open access Nov 2025

Beyond the Headset: A Systematization of Knowledge on Extended Reality Privacy and Security in Healthcare

This SoK surveys 65 peer‐reviewed works published between 2017 and 2024 across leading XR, security, and privacy venues, synthesizing a unified threat taxonomy that spans device, network, user and cloud layers and introduces a quantitative evaluation framework XR-PRISM (Privacy and Risk Impact Scoring Metric).

Nafisa Anjum, M. Mahmud · 4 citations · ⚡2

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.