AgentBoundary: Counterfactual Evaluation of Safety in Tool-Using LLM Agents
This work introduces AgentBound, the first four-way counterfactual generation-and-evaluation framework for tool-using agent safety, and shows that effective agentic alignment requires action decisions to track permission-relevant execution evidence, rather than refusal strength alone.
Tian-Zhuo Yang, Zi-Rui Mi, Yan-Tao Huang et al.
· 1 citation