Continual learning (CL) enables models to acquire new knowledge from sequentially arriving tasks while retaining previously learned knowledge. However, in practical scenarios, task streams collected from untrusted sources may contain backdoor-poisoned samples, posing a critical challenge to the stability, plasticity, a...
This work presents the first bit-flip attack on a VLA: a few gradient-selected flips reduce closed-loop success to $0\%$, while hundreds of random flips are harmless.
Yu-Dong Gao, Ling-Han Chen, Wenhan Wu et al.· 2 citations
Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate increase twentyfold when evaluated on a different dataset. We argue that one contributing f...
PocketAgents is presented, a manifest-driven library of autonomous defense agents that evaluated two agents for the Command and Control and Exfiltration tactics in 18 closed-loop trials of a DarkSide-inspired attack on a small enterprise topology.
Open-weight language models can be rendered unsafe through several parameter-level interventions, yet models with matched harmful compliance can exhibit fundamentally different failure modes. We compare harmful supervised fine-tuning (SFT), harmful reinforcement learning with verifiable rewards (RLVR), and refusal-feat...
Watermarking has emerged as a pivotal solution for content traceability and intellectual property protection in large vision language models (LVLMs). However, vision-agnostic watermarks may introduce visually irrelevant tokens and disrupt visual grounding by enforcing indiscriminate pseudo-random biases. Additionally,...
Traditional distributed backdoor attacks (DBA) in federated learning improve stealthiness by decomposing global triggers into sub-triggers, which however requires more poisoned data to maintian the attck strength and hence increases the exposure risk. To overcome this defect, This paper proposes a novel method, namely...
PiMRef is proposed, the first reference-based solution to detect ever-evolving phishing emails using knowledge-based invariants, targeting the identity-impersonation attacks that characterize spear-phishing.
Ruo-Fan Liu, Yun Lin, Yu-Xin Wang et al.· 4 citations
The results demonstrate that a clean teacher alone is not a sufficient safeguard: poisoned distillation data can produce a strongly backdoored student while maintaining competitive performance on clean images.
Machine unlearning is a promising approach to improve LLM safety by removing unwanted knowledge from the model. However, prevailing gradient-based unlearning methods suffer from issues such as high computational costs, hyperparameter instability, poor sequential unlearning capability, vulnerability to relearning attack...
Aashiq Muhamed, Jacopo Bonato, Mona Diab et al.· 0 citations
Malicious users attempt to replicate commercial models functionally at low cost by training a clone model with query responses. Timely prevention of such model-stealing attacks is challenging, as it requires achieving robust protection, maintaining utility, and ensuring low deployment overhead at the same time. In this...
Jian-Ping Mei, Weibin Zhang, Jie Chen et al.· 0 citations
Large Language Models (LLMs) face significant security risks despite their advanced capabilities. While techniques like Reinforcement Learning with Human Feedback (RLHF) improve ethical alignment, excessive exposure to security-related training data may cause LLMs to overtrust such information, creating new vulnerabili...