Illusion of Alignment: Detecting Hidden Disagreement in Collaborative Dialogue
This work makes IoA detectable by generating diagnostic multiple-choice questions whose divergent answers across participants provide direct behavioral evidence of hidden disagreement, and pairs IoA-Prober-8B with LLM agents improves downstream task performance on BigCodeBench-Hard and HiddenBench.