#artificial intelligence
May 2026
QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents
Evaluating three frontier VLMs in both homogeneous and cross-model adversarial settings, it is found that even the strongest agent hallucinates 15.1% of its verifiable spatial claims and 11.5% of accusations are strictly unsupported.
Ye Yuan, Ruiqi Song, Wei-En Li et al.
· arXiv.org · 2 citations