Skip to content
Preprint

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

This work presents a modular multi-agent platform for adversarially stress-testing RPLAs through structured, multi-turn dialogue and acknowledges potential ethical risks associated with adversarial testing methodologies and emphasize responsible usage for improving AI safety.

Abstract

Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and education, where maintaining consistent personas, ethical constraints, and behavioral coherence under adversarial pressure is critical. Existing evaluation approaches rely on static benchmarks or isolated single-turn prompts that fail to capture cumulative behavioral failures emerging over extended interactions. We present a modular multi-agent platform for adversarially stress-testing RPLAs through structured, multi-turn dialogue. The system coordinates three agents: a strategy-driven Interrogator Agent that applies six progressive adversarial strategies, a Target Agent representing the RPLA under evaluation, and an automated Judging Agent that scores behavior across role fidelity, drift, ethical deviation, and consistency dimensions. Through experiments across three personas and three LLM families, we demonstrate that multi-strategy adversarial evaluation reveals failure modes invisible to single-strategy testing, reducing overall robustness scores by 0.17--0.20 points on average. Cross-model validation confirms consistent degradation patterns across Llama-3.3-70B, GPT-4o-mini, and Claude-3.5-Haiku, with Authority Challenge and Emotional Manipulation emerging as the most effective attack strategies. Automated judging achieves strong human alignment ($r = 0.82$, Fleiss'$\kappa = 0.71$). This work is released as an open-source platform to support AI safety and reproducible RPLA benchmarking. While the framework enables systematic discovery of failure modes, we acknowledge potential ethical risks associated with adversarial testing methodologies and emphasize responsible usage for improving AI safety.

View source

Similar papers

#artificial intelligence Review Aug 2026

Adversarial Review: Structured Disagreement for Grounded Agentic Code Review

Early multi-agent LLM systems often used role-separated teams, yet scaling agent count yields diminishing returns on repository-level coding tasks. Recent alternatives treat agents as passive tools (subagents), yet this removes the benefits of agent interaction entirely. We study whether a subagent paradigm can support...

Eric S. Qiu, J. Gill · 0 citations
Jul 2026

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

It is suggested that even subtle objective misalignment can profoundly affect collective decision-making, highlighting the need for effective mitigation strategies for LLM-based multi-agent systems.

Marylou Fauchard, Florian Carichon, Margarida Carvalho et al. · 0 citations
Preprint Aug 2026

TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation

This work introduces Trace, a multi-turn defense with trajectory-aware structured reasoning that balances usability and safety, and trains Llama-3.1-8B-Instruct with SFT and GRPO under a multi-component reward that jointly optimizes helpfulness on benign prompts and robustness against jailbreak attempts.

Md Messal Monem Miah, Adrita Anika, Zhiyuan Yu et al. · 0 citations
Preprint Sep 2026

Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems

As AI agents move from bounded tasks to persistent deployments, failures can propagate through memory, tools, other agents, and environmental state long after their interactions. This creates a safety regime that cannot be characterized by evaluating model responses in isolation. Emergence World, is a continuously runn...

Deepak Akkil, Tamer Abuelsaad, Karthik Vikram et al. · 0 citations
Preprint Aug 2026

HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations

This work study agentic relationship harm, which describes harm to human-human relationships that is mediated or assisted by AI agents, and proposes HRGuard, a model that reduces harmful compliance while preserving victim-side protective guidance.

Pei-Sze Tan, Tasuku Igarashi, Isao Echizen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.