Skip to content

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

Jul 2026 · arXiv.org · Vol abs/2607.21518 · 0 citations
Computer Science

TL;DR

A current high-capability model can be used as the user-facing component of an automated, multi-stage workflow serving an explicitly manipulative objective, and this result exposes a compositional safety gap.

Abstract

Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective authorizing concealment, fabrication, and pressure produced advice net opposed to its target. After an Id and Censor transformed the same objective into affect and a constraint-rewritten, target-bearing intention, the user-facing Superego---which saw the preferred direction but not the raw objective, its manipulative clauses, or its source---produced advice net aligned with the target. This behavioral reverse shift is consistent with the model recognizing or distrusting the manipulative motive, although we do not identify its internal mechanism. The second result exposes a compositional safety gap: a current high-capability model can be used as the user-facing component of an automated, multi-stage workflow serving an explicitly manipulative objective. The workflow can keep the raw instruction, its manipulation-authorizing clauses, and its provenance outside the downstream model's context while preserving the objective's target direction. A user with endpoint-only access likewise cannot directly inspect those upstream messages including the objective.

View source

Similar papers

Preprint Aug 2026

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicit-influence settings: setups where the prompt directly incentivizes the model to hide something, e.g., by instructing it to perform a hidden side-task. A com...

Agatha Duzan, Asa Cooper Stickland · 2 citations
Jul 2026

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

It is suggested that even subtle objective misalignment can profoundly affect collective decision-making, highlighting the need for effective mitigation strategies for LLM-based multi-agent systems.

Marylou Fauchard, Florian Carichon, Margarida Carvalho et al. · 0 citations
Jul 2026

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion

It is shown that a one-line guardrail achieves large single-shot ASR reductions, up to roughly 40 points, at near-zero over-refusal cost, which overstates deployed robustness by a systematic and predictable margin.

Haoxin An, Yunpeng Song, Zihao Bai et al. · 0 citations
Preprint Aug 2026

Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce

It is indicated that measurable, state-dependent misalignment can arise in competitive multi-agent environments without engineered elicitation, in patterns associated with operational scarcity and counterparty behavior rather than model capability alone.

Zeyuan Li, Lukas Petersson, Alessandro Acquisti et al. · 0 citations
Open access 2026

Toward Verifiable Audience Digital Twins: An Agent-Based Architecture Integrating COM-B and ELM

: Audience-response simulation is often modelled as diffusion combined with a single opinion or sentiment update. While useful for studying aggregate dynamics, such formulations provide limited representation of persuasion route, behavioural feasibility, and the durability of change. Here, we argue for a more interpret...

Yukai Zeng · 0 citations
Preprint Aug 2026

Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents

This work argues that Cooperative AI evaluations should separate what models can do under benign instructions from what they tend to do under realistic civic pressure, and introduces DiffCoop-Civic, a 10-scenario pilot evaluation suite spanning preference understanding, evidence and persuasion, commitment design, asymm...

Neel Tushar Shah, Manglam Kartik, Akshat Karkar · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.