A current high-capability model can be used as the user-facing component of an automated, multi-stage workflow serving an explicitly manipulative objective, and this result exposes a compositional safety gap.
Abstract
Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective authorizing concealment, fabrication, and pressure produced advice net opposed to its target. After an Id and Censor transformed the same objective into affect and a constraint-rewritten, target-bearing intention, the user-facing Superego---which saw the preferred direction but not the raw objective, its manipulative clauses, or its source---produced advice net aligned with the target. This behavioral reverse shift is consistent with the model recognizing or distrusting the manipulative motive, although we do not identify its internal mechanism. The second result exposes a compositional safety gap: a current high-capability model can be used as the user-facing component of an automated, multi-stage workflow serving an explicitly manipulative objective. The workflow can keep the raw instruction, its manipulation-authorizing clauses, and its provenance outside the downstream model's context while preserving the objective's target direction. A user with endpoint-only access likewise cannot directly inspect those upstream messages including the objective.
Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicit-influence settings: setups where the prompt directly incentivizes the model to hide something, e.g., by instructing it to perform a hidden side-task. A com...
It is suggested that even subtle objective misalignment can profoundly affect collective decision-making, highlighting the need for effective mitigation strategies for LLM-based multi-agent systems.
It is shown that a one-line guardrail achieves large single-shot ASR reductions, up to roughly 40 points, at near-zero over-refusal cost, which overstates deployed robustness by a systematic and predictable margin.
Haoxin An, Yunpeng Song, Zihao Bai et al.· arXiv.org· 0 citations
It is indicated that measurable, state-dependent misalignment can arise in competitive multi-agent environments without engineered elicitation, in patterns associated with operational scarcity and counterparty behavior rather than model capability alone.
Zeyuan Li, Lukas Petersson, Alessandro Acquisti et al.· 0 citations
: Audience-response simulation is often modelled as diffusion combined with a single opinion or sentiment update. While useful for studying aggregate dynamics, such formulations provide limited representation of persuasion route, behavioural feasibility, and the durability of change. Here, we argue for a more interpret...
Yukai Zeng· International Conference on...· 0 citations
This work argues that Cooperative AI evaluations should separate what models can do under benign instructions from what they tend to do under realistic civic pressure, and introduces DiffCoop-Civic, a 10-scenario pilot evaluation suite spanning preference understanding, evidence and persuasion, commitment design, asymm...