It is proved that once the fragments look benign in the monitored view, no detector on that view can catch them, however strong it is, and that local safety is not global safety when harm is compositional, and the open problem is finding that representation.
This work builds a working instance on a hierarchical multi-agent system, runs it under benign and attacked conditions across five language models and two task domains, and measures how much of that warning rests on removable surface cues of the attack rather than on its distributed structure.
D. Arias, Dev Prashant Mistry, Ren Wang et al.· arXiv.org· 0 citations
Flip rates are insufficient as a complete measure of open-ended conformity, wrong peers harm open-ended revision, evaluators are not neutral, and anchor calibration is necessary.
Alicia Guerra, Yibo Hu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.