This work presents RedHerring, which inserts certifiably safe decoys that divert verification effort from real vulnerabilities, and shows that its effectiveness does not depend on decoy secrecy.
Kai-Kai Zhang, Zi-Han Zhang, Yu-Chong Xie et al.· 0 citations
Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevolution and validates candidate harnesses or attributed individua...
Zhi-Xiang Zhang, Ze-Sen Liu, Wai-Ip Lai et al.· 0 citations
LLM agents convert model outputs into consequential actions, including communications, code changes, and financial transactions. Developers often trust evidence such as test results and execution logs. We identify a response path integrity gap in Bring Your Own Key configurations used by roughly 88 percent of mainstrea...
Mingyu Luo, Zihan Zhang, Zesen Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.