Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems
This work analyzes the resulting attack-defense setting through a probabilistic model of a target system, its defense mechanism, and the attacker's automated judge, and shows that conventional detect-and-block defenses can allow attacker success rate (ASR) to approach one as the query budget grows.