Skip to content

Category

cybersecurity

1,065 papers

#machine learning Preprint Aug 2026

PAGE: Partition-Aware Gated KV-Cache Eviction

KV-cache eviction can do more than compress. In long-context LLMs, keeping only some cached tokens sometimes matches or exceeds full-cache accuracy, because many redundant prefill tokens otherwise dilute attention away from the tokens that carry the answer. This benefit is not uniform, and evicting the wrong tokens can...

Pankaj Kumar, Subhankar Mishra · 0 citations
#natural language process... Preprint Sep 2026

CASCADE Against Jailbreaks: Combination Across Stages with Controlled Attack-Defense Evaluation

Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine them. Prior empirical studies, fragmented by inconsistent attack-success-rate definitions...

Jia-Le Luo, Eric Han · 1 citation
#natural language process... Preprint Sep 2026

Conformal Privacy Auditing: Calibrated Re-identification Attacks with Statistical Guarantees

Conformal Privacy Auditing is introduced, a distribution-free calibration framework that provides a statistical certificate of re-identification risk for each released document against LLM-empowered adversaries and enables audits of open-source models and proprietary API models in a unified framework.

Shuo Huang, G. Haffari, Xing-Liang Yuan et al. · 0 citations
#machine learning Preprint Open access Sep 2026

The Binary Tree Mechanism is Optimal for Differentially Private Continual Counting

Private continual counting is a fundamental problem in differential privacy: given a binary stream of length $n$, where each $1$ corresponds to the contribution of one individual, the goal is to release all running counts while protecting the privacy of each individual. For fixed privacy parameters, the standard binary...

Konstantina Bairaktari, Markus Engelund Dahl, Kasper Green Larsen · 0 citations
#machine learning Preprint Open access Sep 2026

SteganoBackdoor: Evading Data-Poisoning Defenses via Steganographic Backdoors

Transformer-based models are highly susceptible to backdoor attacks via supervised fine-tuning (SFT). To red-team existing data-poisoning defenses, prior work has increasingly focused on stylized triggers, synthetic artifacts, and token-level perturbations designed to evade detection. However, this trend has shifted th...

Eric Xue, Ruiyi Zhang, Pengtao Xie · 0 citations

OverThink: Slowdown Attacks on Reasoning LLMs

This work evaluates Overthink on proprietary and open-source reasoning models across the FreshQA, SQuAD, and MuSR datasets, and shows that newer generations of RLMs, while showing a drastic increase in per-token cost, also exhibit up to a 2.3x increase in reasoning tokens, leaving them more vulnerable to Overthink atta...

Abhinav Kumar, Jaechul Roh, Ali Naseh et al. · 92 citations · ⚡9
#machine learning Preprint Sep 2026

End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery

The importance of deep neural networks (DNNs) is widely recognized, and the parameters obtained through training are regarded as valuable assets. Recently, attacks that extract these parameters using only oracle queries to a DNN have been actively studied at IACR conferences. The hard-label setting is the most challeng...

Akira Ito, Takayuki Miura, Yosuke Todo · 0 citations
#machine learning Preprint Sep 2026

Watermarkable Multi-Draft Speculative Sampling via Poisson Processes

This work develops a novel multi-draft speculative sampling algorithm based on Poisson processes that maintains both watermark strength and sampling efficiency, and is the first multi-draft, drafter-invariant speculative sampling scheme that maintains both watermark strength and sampling efficiency.

Yan-Xiao Liu, Si-Cheng Wan, Zhan Gao et al. · 1 citation
#machine learning Preprint Open access Sep 2026

ServeGuard: Verifiable, Bounded-Residual Confinement of Operator-Invisible Channels Without Revealing the Certified Read Factor

Third-party adapters for open-weight language models ship as opaque weight matrices; a recipient cannot check whether an adapter hides a backdoor without trusting the publisher or inspecting the weights, the publisher's core asset. For one important class (payloads placed where a safety monitor is structurally blind),...

Dominik Dahlem, Rui Vieira · 0 citations
#machine learning Review Sep 2026

Identifying Security Platform Product Abuse with Machine Learning

This work provides the first study of such a whole-system defense, especially with respect to a deployed and operational capability, and shows an increase in product abuse coverage, a 30% reduction in monthly alerts, and adaptability to changes in malicious actors'behavior.

Shaefer Drew, Michael Brautbar, Paul Knight et al. · 0 citations
#machine learning Preprint Open access Sep 2026

ASGARD: Action-Space Guard for UAV Resilience via Reinforcement Learning

Reinforcement learning (RL) controllers have been recently adopted for Unmanned Aerial Vehicles (UAV) navigation and control. However, they are susceptible to action-space attacks that overwrite the action commands after the policy generates them and before the actuators execute them. While most existing defenses targe...

Mohsen Salehi, Karthik Pattabiraman · 0 citations
#artificial intelligence Preprint Sep 2026

TERMon: Detecting Persistent Behavioral Threats in Edge AI via Hardware-Native Ternary Runtime Monitor

Edge AI accelerators are increasingly deployed in safety-critical environments, where model outputs may control physical actuators, make access-control decisions, or trigger alarms. In these settings, runtime failures often remain undetected because model corruption, distribution shift, and adversarial inputs can still...

Arish Sateesan, Edlira Dushku · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.