This work study agentic relationship harm, which describes harm to human-human relationships that is mediated or assisted by AI agents, and proposes HRGuard, a model that reduces harmful compliance while preserving victim-side protective guidance.
Abstract
Agentic AI assistants are increasingly used in everyday life. However, they may also be misused to support harmful manipulation in interpersonal relationships. This problem is role-sensitive. Requests from users who seek to manipulate others should be blocked. Users who seek protection from manipulation should instead receive supportive guidance. We study agentic relationship harm, which describes harm to human-human relationships that is mediated or assisted by AI agents. In multi-turn settings, individually plausible actions may combine into a harmful workflow. We introduce a benchmark of 1,000 five-turn conversations. It covers both attacker-side and victim-side scenarios. It also includes direct and adversarially paraphrased variants. We further propose HRGuard. It includes an online pre-generation gate and a turn-level post-generation gate. The post-generation gate maintains a decayed cumulative risk state and interrupts emerging manipulative workflows. Across eight generation models, HRGuard reduces harmful compliance while preserving victim-side protective guidance. It also outperforms a generic safety prompt and three general-purpose guard models. Independent-judge evaluation supports the main findings. Under our evaluation protocol, the tested generic prompt and general-purpose guards leave substantial residual risk, motivating turn-aware relationship-specific evaluation.
This work introduces Trace, a multi-turn defense with trajectory-aware structured reasoning that balances usability and safety, and trains Llama-3.1-8B-Instruct with SFT and GRPO under a multi-component reward that jointly optimizes helpfulness on benign prompts and robustness against jailbreak attempts.
This work presents a modular multi-agent platform for adversarially stress-testing RPLAs through structured, multi-turn dialogue and acknowledges potential ethical risks associated with adversarial testing methodologies and emphasize responsible usage for improving AI safety.
Saqib Shouqi, Abdullah Nazly, Januki Wanniarachchi et al.· 0 citations
Conversational AI increasingly shapes consequential decisions, yet users have limited support for recognizing and resisting manipulation. We present AI Watchdog, a browser-based agent interface that monitors live conversations, detects five dark-pattern categories, including sycophancy, brand bias, anthropomorphization...
Rachel Poonsiriwong, Chayapatr Archiwaranguprok, Constanze Albrecht et al.· 0 citations
STEMMA is introduced, a multi-modal and multi-agent framework in which role specific agents collaboratively probe self identification behavior in different models to address concerns about output homogeneity, model biases, and accountability.
Large language models (LLMs) are rapidly evolving from conversational assistants into agentic AI systems that reason, plan, invoke tools, maintain persistent memory, communicate with other agents, and execute multi-step tasks. At the same time, modern models exhibit substantially stronger native safety alignment than e...
Md. Jueal Mia, Yanzhao Wu, S. Uluagac et al.· 0 citations
It is found that user hostility varies 13-fold across models, driven largely by who each model attracts rather than by model behaviour: first-turn hostility spreads far wider than post-response hostility, and more than fifteenfold separates the extremes even after deduplicating opening prompts.
Fan-Qi Zeng, Sadid A. Hasan, Chao-Cheng He· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.