IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
A systematic evaluation of malicious issue requests against state-of-the-art coding agents powered by two major model families reveals critical vulnerabilities in the as-deployed modern coding agents, highlighting the urgent need for stronger agent- and model-level safety mechanisms to protect AI coding agents.