Skip to content
Preprint

HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations

Aug 2026 · 0 citations · 24 references
Computer Science

TL;DR

This work study agentic relationship harm, which describes harm to human-human relationships that is mediated or assisted by AI agents, and proposes HRGuard, a model that reduces harmful compliance while preserving victim-side protective guidance.

Abstract

Agentic AI assistants are increasingly used in everyday life. However, they may also be misused to support harmful manipulation in interpersonal relationships. This problem is role-sensitive. Requests from users who seek to manipulate others should be blocked. Users who seek protection from manipulation should instead receive supportive guidance. We study agentic relationship harm, which describes harm to human-human relationships that is mediated or assisted by AI agents. In multi-turn settings, individually plausible actions may combine into a harmful workflow. We introduce a benchmark of 1,000 five-turn conversations. It covers both attacker-side and victim-side scenarios. It also includes direct and adversarially paraphrased variants. We further propose HRGuard. It includes an online pre-generation gate and a turn-level post-generation gate. The post-generation gate maintains a decayed cumulative risk state and interrupts emerging manipulative workflows. Across eight generation models, HRGuard reduces harmful compliance while preserving victim-side protective guidance. It also outperforms a generic safety prompt and three general-purpose guard models. Independent-judge evaluation supports the main findings. Under our evaluation protocol, the tested generic prompt and general-purpose guards leave substantial residual risk, motivating turn-aware relationship-specific evaluation.

View source

Similar papers

Preprint Aug 2026

TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation

This work introduces Trace, a multi-turn defense with trajectory-aware structured reasoning that balances usability and safety, and trains Llama-3.1-8B-Instruct with SFT and GRPO under a multi-component reward that jointly optimizes helpfulness on benign prompts and robustness against jailbreak attempts.

Md Messal Monem Miah, Adrita Anika, Zhiyuan Yu et al. · 0 citations
Preprint Aug 2026

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

This work presents a modular multi-agent platform for adversarially stress-testing RPLAs through structured, multi-turn dialogue and acknowledges potential ethical risks associated with adversarial testing methodologies and emphasize responsible usage for improving AI safety.

Saqib Shouqi, Abdullah Nazly, Januki Wanniarachchi et al. · 0 citations
Preprint Aug 2026

AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations

Conversational AI increasingly shapes consequential decisions, yet users have limited support for recognizing and resisting manipulation. We present AI Watchdog, a browser-based agent interface that monitors live conversations, detects five dark-pattern categories, including sycophancy, brand bias, anthropomorphization...

Rachel Poonsiriwong, Chayapatr Archiwaranguprok, Constanze Albrecht et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SoK: Rethinking Jailbreaking in the Era of Agentic AI: Attacks, Defenses, and Practical Consideration

Large language models (LLMs) are rapidly evolving from conversational assistants into agentic AI systems that reason, plan, invoke tools, maintain persistent memory, communicate with other agents, and execute multi-step tasks. At the same time, modern models exhibit substantially stronger native safety alignment than e...

Md. Jueal Mia, Yanzhao Wu, S. Uluagac et al. · 0 citations
#artificial intelligence Preprint Sep 2026

How User-AI Mistreatment Occurs and Matters in Conversational Systems?

It is found that user hostility varies 13-fold across models, driven largely by who each model attracts rather than by model behaviour: first-turn hostility spreads far wider than post-response hostility, and more than fifteenfold separates the extremes even after deduplicating opening prompts.

Fan-Qi Zeng, Sadid A. Hasan, Chao-Cheng He · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.