The proposed Multi-turn Multimodal Interactive Dialogue (MMID) Benchmark enables comprehensive evaluation of the Perception, Memorization, and Reasoning abilities of MLLMs, and reveals MLLMs perform well with text-based input but degrade with images, requiring improved leverage fine-grained visual cues.
Seulgi Kim, Juoh Sun, Sumin Kim et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.