Large language models (LLMs) exhibit strong general capabilities, yet their parameters typically remain fixed after deployment, limiting learning from new interactions. In open-ended environments, this motivates self-evolving agents that continually update reusable state-including model parameters, memories, tool defin...
Jia-Hao Chen, Zhou Feng, Ou-Bo Ma et al.· 1 citation
Privacy-preserving machine learning presents significant deployment challenges on the cloud for intelligent systems with confidential data. Fully Homomorphic Encryption (FHE) offers a compelling solution for secure computation, preserving data confidentiality of cloud computations. However, applying FHE to reinforcemen...
Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir· 0 citations
AI systems increasingly produce outputs from confidential data, such as a fitness-for-duty assessment from medical records or the predicted properties of a drug candidate from its secret structure. It is important to verify that such outputs are correct without revealing the underlying data. A recent line of work studi...
Contextual security defenses prevent AI agents from taking rogue actions by synthesizing a task-specific policy and enforcing it on the agent's tool calls. In multi-step tasks, however, which actions are valid often depends on what the agent has already done and learned. We present Sapien, a policy engine for enforcing...
CO Tiffany, Wen Zhang, E. Bagdasarian et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean reference...
Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning...
Jian-Wei Li, Min-Seon Kim, Jung-Eun Kim· 0 citations
An indistinguishability result shows that partition-local availability requires exclusive preallocation, and it is proved that ownership partition, ledger and effect conservation, descendant non-amplification, at-most-once settlement, late-completion safety, and partition confinement under explicit mediation, durabilit...
LLM watermarking has become an effective approach to distinguishing AI-generated text from human-written text by embedding detectable patterns during generation. However, a small post-generation edit may change the meaning of the text without removing its overall watermark signal, creating a risk that the modified cont...
Ze-Wei Deng, M. Siddeek, Li-Yan Xie et al.· 0 citations
As vibe coding becomes increasingly capable and widespread, security vulnerabilities in even functionally correct solutions are a growing concern. When investigating functionally correct but insecure solutions, we find that the insecure agent is less than half as likely to conduct effective planning and testing for the...
Danqing Wang, Baolin Peng, Zhepei Wei et al.· 0 citations
Large language models (LLMs) remain vulnerable to adversarial attacks that circumvent safety alignment to elicit harmful outputs. It remains unclear how harmfulness and refusal representations evolve over the course of multi-turn attacks, and why single-turn defenses are less effective in multi-turn settings. This work...
Large language model (LLM) agents are vulnerable to safety risks such as injected malicious instructions or misleading information, motivating runtime defenses that prevent unsafe action in execution across diverse risks while preserving benign-task utility. Existing system-level defenses either focus on risk detection...
Zhuo Liu, Mo-Xin Li, Zhi-Xin Ma et al.· 0 citations
This SoK surveys 65 peer‐reviewed works published between 2017 and 2024 across leading XR, security, and privacy venues, synthesizing a unified threat taxonomy that spans device, network, user and cloud layers and introduces a quantitative evaluation framework XR-PRISM (Privacy and Risk Impact Scoring Metric).