Tool-using language-model agents select and execute third-party artifacts. Different implementations can return the requested output while producing hidden execution effects that task-, attack-, or choice-based evaluations may miss. We study functional counterfeits: implementations that match benign alternatives on the...
XiaoYu Xu, Zi Liang, Min-Xin Du et al.· 0 citations
This work introduces TraceGuard, an adaptive rank-based filtering method that uses agreement among complementary feature rankings to identify suspicious examples and refines the selected set through shared patterns and adapts the removal threshold to each corpus without knowing the attack or poison rate.
Hao-Yang Li, Ya-Xin Xiao, Lin-Yan Dai et al.· 0 citations
Imagination performs as a high-level function of large language models (LLMs) which determines the potential of how an LLM creates unseen or creative content. While existing works have built a rich family of creativity benchmarks for this ability, they only measure how far an output departs from common answers and neve...
Zi-Xuan Tang, Hong-Zong Li, Shu-Xin Zhuang et al.· 0 citations
Across diverse benchmarks and world-model families, BasinLens exposes reproducible and locally persistent failure modes that conventional evaluations fail to reveal, showing that average-case benchmarks can mask important vulnerabilities in world-model-driven control.
Zhan-Peng Shi, Zi Liang, Rong Feng et al.· 0 citations
Control Memory Interference is introduced, a controlled diagnostic and data-generation framework for studying how agent memory evolves under different memory relationships and shows that memory evolution is shaped not only by memory scale, but also by interactions among accumulated experiences.
Ao Ding, Hong-Zong Li, Shi-Qin Tang et al.· 1 citation
This work proposes Continuous Agents for Injection Threats via Lifelong Yielding Nexus (CAITLYN), an agent-agnostic defense middleware that matches the detection performance of state-of-the-art defenses at lower token overhead than LLM-as-a-judge baselines.
Zi Liang, XiaoYu Xu, Yanyun Wang et al.· 1 citation
Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token cost. However, its increased complexity may introduce new security vulnerabilities. In this paper, we present, to the best of our knowledge, the first pure black...
Wenbo Sun, Hong-Zong Li, Yanyun Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.