Compliance without coherence: fluent failure and the ethics of alignment evaluation in multi-agent language models
This essay argues that the deployed monitoring layer—distinct from red-teaming, agent benchmarks, long-context consistency testing, and interpretability research, which it complements—has a principled blind spot, and argues that deployed alignment evaluation should be epistemically plural, complementing behavioral monitors with structural audits of an agent’s discourse.