KV-cache eviction can do more than compress. In long-context LLMs, keeping only some cached tokens sometimes matches or exceeds full-cache accuracy, because many redundant prefill tokens otherwise dilute attention away from the tokens that carry the answer. This benefit is not uniform, and evicting the wrong tokens can...
Defenses against jailbreak attacks on Large Language Models (LLMs) operate at different pipeline stages, such as input modification or output guard, but it remains unclear which defenses to deploy at each stage and how to combine them. Prior empirical studies, fragmented by inconsistent attack-success-rate definitions...
Conformal Privacy Auditing is introduced, a distribution-free calibration framework that provides a statistical certificate of re-identification risk for each released document against LLM-empowered adversaries and enables audits of open-source models and proprietary API models in a unified framework.
Shuo Huang, G. Haffari, Xing-Liang Yuan et al.· 0 citations
Private continual counting is a fundamental problem in differential privacy: given a binary stream of length $n$, where each $1$ corresponds to the contribution of one individual, the goal is to release all running counts while protecting the privacy of each individual. For fixed privacy parameters, the standard binary...
Konstantina Bairaktari, Markus Engelund Dahl, Kasper Green Larsen· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Transformer-based models are highly susceptible to backdoor attacks via supervised fine-tuning (SFT). To red-team existing data-poisoning defenses, prior work has increasingly focused on stylized triggers, synthetic artifacts, and token-level perturbations designed to evade detection. However, this trend has shifted th...
This work evaluates Overthink on proprietary and open-source reasoning models across the FreshQA, SQuAD, and MuSR datasets, and shows that newer generations of RLMs, while showing a drastic increase in per-token cost, also exhibit up to a 2.3x increase in reasoning tokens, leaving them more vulnerable to Overthink atta...
Abhinav Kumar, Jaechul Roh, Ali Naseh et al.· arXiv.org· 92 citations· ⚡9
The importance of deep neural networks (DNNs) is widely recognized, and the parameters obtained through training are regarded as valuable assets. Recently, attacks that extract these parameters using only oracle queries to a DNN have been actively studied at IACR conferences. The hard-label setting is the most challeng...
This work develops a novel multi-draft speculative sampling algorithm based on Poisson processes that maintains both watermark strength and sampling efficiency, and is the first multi-draft, drafter-invariant speculative sampling scheme that maintains both watermark strength and sampling efficiency.
Yan-Xiao Liu, Si-Cheng Wan, Zhan Gao et al.· 1 citation
Third-party adapters for open-weight language models ship as opaque weight matrices; a recipient cannot check whether an adapter hides a backdoor without trusting the publisher or inspecting the weights, the publisher's core asset. For one important class (payloads placed where a safety monitor is structurally blind),...
This work provides the first study of such a whole-system defense, especially with respect to a deployed and operational capability, and shows an increase in product abuse coverage, a 30% reduction in monthly alerts, and adaptability to changes in malicious actors'behavior.
Shaefer Drew, Michael Brautbar, Paul Knight et al.· 0 citations
Reinforcement learning (RL) controllers have been recently adopted for Unmanned Aerial Vehicles (UAV) navigation and control. However, they are susceptible to action-space attacks that overwrite the action commands after the policy generates them and before the actuators execute them. While most existing defenses targe...
Edge AI accelerators are increasingly deployed in safety-critical environments, where model outputs may control physical actuators, make access-control decisions, or trigger alarms. In these settings, runtime failures often remain undetected because model corruption, distribution shift, and adversarial inputs can still...