HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations
This work study agentic relationship harm, which describes harm to human-human relationships that is mediated or assisted by AI agents, and proposes HRGuard, a model that reduces harmful compliance while preserving victim-side protective guidance.