Skip to content

Category

cybersecurity

1,065 papers

#machine learning Preprint Sep 2026

Robust Dynamic Expansion for Continual Learning under Backdoor Attacks via Purification and Selective Recovery

Continual learning (CL) enables models to acquire new knowledge from sequentially arriving tasks while retaining previously learned knowledge. However, in practical scenarios, task streams collected from untrusted sources may contain backdoor-poisoned samples, posing a critical challenge to the stability, plasticity, a...

Ke-Yu Lin, Fei Ye, Qi-He Liu et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate increase twentyfold when evaluated on a different dataset. We argue that one contributing f...

Daniyal Kabir Dar, Arun Ross · 0 citations

PocketAgents: A Manifest-Driven Library of Autonomous Defense Agents

PocketAgents is presented, a manifest-driven library of autonomous defense agents that evaluated two agents for the Command and Control and Exfiltration tactics in 18 closed-loop trials of a DarkSide-inspired attack on a small enterprise topology.

Sidnei Barbieri, Ágney Lopes Roth Ferraz, L. Pereira · 1 citation
#artificial intelligence Preprint Open access Sep 2026

Behind Harmful Compliance: Behavioral and Mechanistic Divergence Across LLM Jailbreaks

Open-weight language models can be rendered unsafe through several parameter-level interventions, yet models with matched harmful compliance can exhibit fundamentally different failure modes. We compare harmful supervised fine-tuning (SFT), harmful reinforcement learning with verifiable rewards (RLVR), and refusal-feat...

Md Rysul Kabir, Zoran Tiganj · 0 citations
#artificial intelligence Preprint Open access Sep 2026

AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language Models

Watermarking has emerged as a pivotal solution for content traceability and intellectual property protection in large vision language models (LVLMs). However, vision-agnostic watermarks may introduce visually irrelevant tokens and disrupt visual grounding by enforcing indiscriminate pseudo-random biases. Additionally,...

Yue Li, Xin Yi, Dongsheng Shi et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Unveiling Hidden Threats: Using Fractal Triggers to Boost Stealthiness of Distributed Backdoor Attacks in Federated Learning

Traditional distributed backdoor attacks (DBA) in federated learning improve stealthiness by decomposing global triggers into sub-triggers, which however requires more poisoned data to maintian the attck strength and hence increases the exposure risk. To overcome this defect, This paper proposes a novel method, namely...

Jian Wang, Hong Shen, Chan-Tong Lam · 0 citations
#artificial intelligence Preprint Apr 2025

How to Backdoor Image Knowledge Distillation

The results demonstrate that a clean teacher alone is not a sufficient safeguard: poisoned distillation data can produce a strongly backdoored student while maintaining competitive performance on clean images.

Qian Ma, Chen Wu, P. Mitra et al. · 3 citations
#artificial intelligence Preprint Open access Sep 2026

SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs

Machine unlearning is a promising approach to improve LLM safety by removing unwanted knowledge from the model. However, prevailing gradient-based unlearning methods suffer from issues such as high computational costs, hyperparameter instability, poor sequential unlearning capability, vulnerability to relearning attack...

Aashiq Muhamed, Jacopo Bonato, Mona Diab et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

D-ADD: An Effective Plug-In for Defending Against Model Stealing

Malicious users attempt to replicate commercial models functionally at low cost by training a clone model with query responses. Timely prevention of such model-stealing attacks is challenging, as it requires achieving robust protection, maintaining utility, and ensuring low deployment overhead at the same time. In this...

Jian-Ping Mei, Weibin Zhang, Jie Chen et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

FATS: A Prompt Injection Attack Utilizing Feign Security Agents with Deceptive Few-shots Learning

Large Language Models (LLMs) face significant security risks despite their advanced capabilities. While techniques like Reinforcement Learning with Human Feedback (RLHF) improve ethical alignment, excessive exposure to security-related training data may cause LLMs to overtrust such information, creating new vulnerabili...

Yupeng Ren, Jiangtao Chen, Rui Zhang · 0 citations

From tech blogs

See all →
Google DeepMind Blog Jul 17, 2026

Introducing Gemini 3.5 Flash Cyber

Google introduces Gemini 3.5 Flash Cyber, a lightweight cybersecurity model to find and patch vulnerabilities.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.