Do User-Authored Permission Policies Improve Protection Against AI Agent Overreach?
A mechanism in which a language model maps actions to plain-language consequence categories with user-authored"allow","ask", or"never"rules is examined, and a gap between preference and commitment is revealed: repeatedly choosing "ask" preserves case-by-case choice but prevents a standing policy from settling decisions in advance.