Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions
It is shown that hallucinations often accumulate over long contexts, through self-reinforcing dialogue history, and models are particularly vulnerable to questions requiring premise rejection or refusal, which highlights CEDI as a step toward realistic, systematic, and ecologically valid assessments of MLLMs'capabiliti...