Skip to content

Category

cybersecurity

1,032 papers

#natural language process... Preprint Open access Oct 2026

Sensitive-Topic Leakage Through LLM Routing Metadata: Measurement and Mitigation

LLM routers pick a cheap or expensive model per request by its content, and many gateways and some cloud platforms can log that choice with content logging off. We measure this privacy channel beyond token counts, accounting for noisy labels and repeated prompts. We run pre-registered studies on 1.7 million real reques...

Teng-Ruei Chen · 0 citations
#machine learning Preprint Open access Oct 2026

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks

Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful behavior is learned through fine-tuning rather than elicited by jailbreaking the model. Yet, pretrained LLMs already encode substantial harmful...

Kevin Kuo, Virginia Smith, Chhavi Yadav · 0 citations
#machine learning Preprint Open access Oct 2026

Estimating Model-Level Membership Inference Vulnerability Without Reference Models

Membership inference attacks (MIAs) have emerged as the standard tool for evaluating the privacy risks of AI models. However, state-of-the-art attacks require training numerous, often computationally expensive, reference models, limiting their practicality. We present a novel approach for estimating model-level vulnera...

Euodia Dodd, Nata\v{s}a Kr\v{c}o, Igor Shilov et al. · 0 citations
#machine learning Preprint Open access Oct 2026

The Utility and Complexity of in- and out-of-Distribution Machine Unlearning

Machine unlearning, the process of selectively removing data from trained models, is increasingly crucial for addressing privacy concerns and knowledge gaps post-deployment. Despite this importance, existing approaches are often heuristic and lack formal guarantees. In this paper, we analyze the fundamental utility, ti...

Youssef Allouah, Joshua Kazdan, Rachid Guerraoui et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Pareto-optimal quantum kernel selection for unsupervised anomaly detection on real malware beaconing data

Quantum kernel methods are leading candidates for a practical quantum advantage in machine learning, but assessing that potential requires two quantities usually reported separately: how well a kernel performs on the task, and how far its geometry departs from the classical kernels available for the same problem. We in...

Boaz Micah, Nadia Milazzo, Maissa Beji et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs

Token-Pruning accelerates Vision-Language Models by removing redundant visual tokens, yet its safety implications remain underexplored. In this work, we present the first comprehensive safety evaluation of Token-Pruning mechanisms and find that: most pruning strategies significantly degrade safety as pruning ratios inc...

Shuailong Wang, Xinyu Lyu, Shengming Yuan et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Constitution-Guided Watermarking

Watermarking enables language model providers to identify text generated by their models. However, its desired properties can conflict (\ie~stronger watermark signals can degrade text quality), while designs that resist editing may also facilitate forgery. Providers address these trade-offs by choosing configurations t...

Toluwani Aremu, Samuele Poppi, Nils Lukas · 0 citations
#machine learning Preprint Open access Oct 2026

Ask the Expert: LLM-Guided Reinforcement Learning for Autonomous Cyber Defense

Policy-based reinforcement learning (RL) approaches have produced promising results for autonomous cyber defense; however, they are sample-inefficient in settings where defenders must respond under delayed, partial observations with actions from large action spaces. While large language models (LLMs) may reason semanti...

Fernando Martinez, Abhishek Satyam, Tao Li et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Adversarial RL for Port-Scan Evasion: Attacker Feature Visibility in Edge-Deployed IDS

Machine learning-based intrusion detection systems (IDS) are increasingly used in resource-constrained Internet of Things (IoT) environments, yet their robustness is often evaluated against static attacks rather than adversaries that adapt to detection feedback. This paper investigates adaptive port-scan evasion agains...

Logan Andrew North, Priya Sanjay Kaluskar, Shasi Kumar Ramachandran Prabhu et al. · 0 citations
#machine learning Preprint Open access Oct 2026

PatchBench: Measuring Collateral Damage in Activation Patching

An LLM safety patch can pass a benchmark while still being a poor repair. This risk is especially acute for jailbreak repairs, where the goal is to correct a specific unsafe behaviour without changing unrelated behaviours. A patch may block exact evaluation prompts yet fail on close harmful variants, or suppress harmfu...

Alexi Canesse, Mathis Le Bail, Ma\"el Jenny et al. · 0 citations
#machine learning Preprint Open access Oct 2026

PairAudit: Guiding Human Review with Graph Tokens under Distribution Shift

Intrusion detectors can confidently misclassify attacks that were not seen during training. Human review can correct these errors, but only a limited number of cases can be checked. Uncertainty-based review may overlook confident errors, while anomaly scores alone do not show whether changing the review plan will corre...

Jiran Tao, Binyan Jiang · 0 citations
#machine learning Preprint Open access Oct 2026

Efficient Provably Private Classification with a Tabular Foundation Model

Tabular data underpin prediction and decision-making in medicine, finance, government and science, but often contain sensitive individual-level information, creating a need for accurate prediction while preserving privacy. Traditional private learning provides formal privacy guarantees, but requires slow dataset-specif...

Talal Alrawajfeh, Cristiana Diaconu, Ossi R\"ais\"a et al. · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.