Skip to content

Category

cybersecurity

1,032 papers

#artificial intelligence Preprint Open access Oct 2026

Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents

Enterprise AI agents that share a memory store face two unaddressed risks: sensitive data can leak through legitimately computed results the requester could not derive, and departments can silently compute a same-named key performance indicator (KPI) through conflicting logic. Existing agent-memory systems (e.g., MemGP...

Venkata M Sangaraju, Sudhir Vissa · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Jailbreaking Open-Weight LLMs via Random Embedding Perturbations

While open-weight models have enjoyed steady progress in capabilities and wide adoption across multiple domains, their safety remains an important concern. One key feature is the ability to refuse or deflect harmful, malicious, or insensitive prompts. In this paper, we expose safety vulnerabilities across six common op...

Abhinav Sudhakar Dubey (University of California Santa Cruz), Scott Sirri (University of California Santa Cruz), Vaggos Chatziafratis (University of California Santa Cruz) et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Towards a Unified Misuse Monitoring Benchmark

LLM agents increasingly act in multi-actor environments, exposing them to misuse from multiple sources: decomposition attacks, where a harmful request is split into innocuous sub-requests, and prompt injection attacks, where a compromised tool delivers a malicious instruction. Existing evaluations treat these threats s...

Aniruddh Pramod, James Oldfield, Adel Bibi · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Which Image Property Carries the Jailbreak? A Controlled Dissection of Image-to-Text Jailbreaks

Image-to-text jailbreaks place harmful intent in text, image content, or the relationship between them. We examine image-side factors across four published attack families on a 313-prompt StrongREJECT slice, using five multimodal models and an additional appendix evaluation of InternVL3.5-8B. The harmful instruction is...

Boyuan Chen, Yehia Dawoud, Hailemariam Mersha et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Where Does the Audio Jailbreak Live? A Controlled Frequency-Depth Audit of AdvWave-P on Qwen2-Audio

We audit frequency and decoder-depth claims for AdvWave-P, an additive audio jailbreak, on Qwen2-Audio. The protocol masks frequency components of the perturbation in the short-time Fourier transform (STFT) domain and measures attack success and audio-span representations. On 520 AdvBench prompts, the primary judge lab...

Boyuan Chen, Minseok Kim, Sohaila Abdulsattar et al. · 0 citations
#machine learning Preprint Open access Oct 2026

On the Intrinsic Limited Robustness of Latent-Based Watermarking

Existing latent-based watermarking methods for diffusion models have overestimated their robustness to image distortions, including geometric transformations such as rotation, scaling, and translation (RST). Moreover, this paradigm of watermarking approaches may suffer from inherent limitations arising from the domain...

Cheng-Han Yeh, Kuan-chun Yu, Cheng-Chang Tsai et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Adaptive Model Inversion Attacks Generalize a Privacy-Robustness Tradeoff

In this paper, we show that standard evaluations of high-resolution Model Inversion Attacks (MIAs) significantly underestimate training-data privacy leakage. State-of-the-art privacy defenses, standard training techniques such as MixUp and Adversarial Training, and undefended models all leak training images at rates 1....

Shailen Smith, Rasmus Torp, Adam Breuer · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates

Safety auditing of post-trained large language models typically relies on model behavior, requiring model execution and depending on the coverage of available evaluations. This work asks a different question: Do the target behaviors optimized during supervised fine-tuning (SFT) leave readable evidence directly in check...

Ziqun Bao, Xinyu Zhang, Yuchen Shao et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Weight Oracles: Reading Neural Network Weights with Language Models

Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known inputs to find hidden capabilities such as backdoors. We propose Weight Oracles, fine-tuned language models that diagnose properties of a target network by reading i...

Krishna Kabra, Constantin Venhoff, Christian Schroeder de Witt · 0 citations
#machine learning Preprint Open access Oct 2026

ATLAS-AL: Adaptive Trust-Region for Latent Adversarial Searches via Active Learning

Security evaluation of learning-based systems requires more than just testing the system against a fixed collection of attacks. It requires adaptive mechanisms that can efficiently discover \textit{sets} of inputs that induce model failure. We introduce ATLAS (Adaptive Trust-Regions for Latent Adversarial Searches), wh...

Marsalis Gibson, Claire Tomlin, Shankar Sastry · 0 citations
#machine learning Preprint Open access Oct 2026

Reward-Driven Learning under Prompt-Level Differential Privacy

Reinforcement learning with verifiable rewards (RLVR) trains a language model on problems that may themselves be confidential, and the trained model can reveal which problems it saw. We study RLVR under prompt-level differential privacy: the released weights must be ({\epsilon},{\delta})-differentially private with res...

Jiachen Zhao, Antonia Januszewicz, Taeho Jung · 0 citations
#natural language process... Preprint Open access Oct 2026

Cross-Lingual Summarization as a Black-Box Watermark Removal Attack

Watermarking has been proposed as a lightweight mechanism to identify AI-generated text, with schemes typically relying on perturbations to token distributions. While prior work shows that paraphrasing can weaken such signals, these attacks remain partially detectable or degrade text quality. We demonstrate that cross-...

Gokul Ganesan · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.