#large language models
Jul 2026
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation
A current high-capability model can be used as the user-facing component of an automated, multi-stage workflow serving an explicitly manipulative objective, and this result exposes a compositional safety gap.
Lin-Jun Li
· arXiv.org · 0 citations