Large language models (LLMs) exhibit strong general capabilities, yet their parameters typically remain fixed after deployment, limiting learning from new interactions. In open-ended environments, this motivates self-evolving agents that continually update reusable state-including model parameters, memories, tool defin...
Jia-Hao Chen, Zhou Feng, Ou-Bo Ma et al.· 0 citations
CIPR (Coding In Poisoned Repos), the first benchmark that systematically varies PLCs in poisoned real-world repositories, is introduced and highlights that coding agent vulnerability is not a static property, but a dynamic outcome shaped by everyday user configurations.
Fu-Kang Zhu, Bin-Bin Zhao, Rui-Xiao Lin et al.· 0 citations
AEGIS (Adaptive Ensemble Guard for Injection Shielding) extracts instruction-sensitive projectors to identify malicious instructions and leverages a Unified Multi-Layer Consensus mechanism that aggregates topologically distinct signals across the network depth.
Jia-Hao Chen, Ruiping Yin, Xin-Feng Li et al.· 1 citation· ⚡1
Unsafe Semantic Distillation is proposed, which aligns adversarial perturbations with distributional representations of unsafe content rather than prompt-specific instances, and achieves 84% attack success rates, outperforming existing methods and exposing fundamental vulnerabilities in current multimodal safety archit...
Shuo Shi, Ruiping Yin, Na-En Xu et al.· Proceedings of the 32nd ACM...· 1 citation
MTGuard is proposed, a hybrid analysis-based defense framework designed to safeguard the use of MCP tools in LLM agents by leveraging lifecycle-aware static-dynamic co-analysis and effectively mitigates multiple categories of harmful tool use across different LLM agents while maintaining performance on benign user task...
Ping He, Yuexiang Xie, Yaliang Li et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.