Skip to content
Book Open access

MVSCbench: A Benchmark for Symbolic and Cultural Understanding in Music Videos

Jul 2026 · Creativity & Cognition · 0 citations · 20 references
Computer Science

Abstract

This study introduces MVSCBench, the benchmark designed to evaluate multimodal AI models’ ability to interpret symbolic and cultural meanings in music videos. Using K-pop music video clips, we organize music video understanding into three stages: surface perception, symbolic interpretation, and cultural grounding. Based on this framework, we construct a dataset of 293 video clips and 3,516 annotations, covering space, artist, character symbolism, object symbolism, literary symbolism, social-historical symbolism, and fandom cultural symbols. The results show that current multimodal AI models perform well on surface perception tasks, but still face clear limitations in symbolic interpretation and cultural grounding. In particular, the models often struggle to correctly identify relevant visual cues and frequently produce hallucinated interpretations in categories involving deeper meanings. MVSCBench provides a new benchmark for evaluating symbolic and cultural understanding in music videos and offers a useful framework for future research.

Read PDF

Similar papers

Preprint Aug 2026

CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

Results show that although current models achieve strong semantic adherence and visual quality, they often fail to faithfully capture fine-grained cultural details, particularly for underrepresented regions, rituals, and multimodal cultural cues.

Xianjing Han, Yuhan Su, Yang Deng et al. · 2 citations
#artificial intelligence Review Aug 2026

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

An overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation is presented, covering spoken visual question answering and image-grounded hallucination detection in English and Modern Standard Arabic, and CRAI-Bench, evaluating the cultural accuracy of text-to-image generation.

Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti et al. · 0 citations
#small language model Preprint Aug 2026

PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding

PUMA (Polish Unified Multimodal Assessment) is proposed, a novel benchmark of 900 hand-crafted tasks designed to probe the limits of multimodal models in the Polish cultural and linguistic context and open-source the evaluation framework to advance localized multimodal AI research.

Slawomir Dadas, Michał Perełkiewicz, Rafal Poswiata et al. · 0 citations
Preprint Aug 2026

Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of cultural concepts. We decompose this into three abilities: naming what a concept symbolizes, visually recognizing it on video, and locating its sub-events in time. Existing video-cultural benchmarks tend to test what is seen, collapsing these three abilities into a single score that hides the bottleneck. We introduce the Cultural Moment Benchmark (CMB): 306 expert-curated concepts from seven countries in Southeast Asia across five categories. We evaluate each concept through three stages, one per ability. Given a description, Stage 1 (S1) selects from four candidate concept names, Stage 2 (S2) selects from four candidate video moments, and Stage 3 (S3) predicts the start and end times of the moment in a video. To keep each stage focused on a distinct ability, we use three design choices: semantic-similarity distractors (S1, S2), unlabeled video moments (S2), and free-form localization on a different example video (S3). Across six vision-language models, failure modes vary by ability and modality. i) Even the strongest closed-source models score below 30% when all three stages must be correct; ii) The three abilities do not fully cascade: naming a concept correctly helps half the models recognize it on video, but recognizing it has little effect on locating the sub-event in time; iii) Audio is complementary, redundant, or distracting depending on the concept, more often distracting in non-Latin-script countries; removing both audio and subtitles hurts Games and Music the most. Our 14-rater human study shows that even Expert raters score below chance on concepts from a neighboring country, indicating that CMB requires country-specific cultural knowledge. CMB acts as a diagnostic harness, attributing failures to a specific ability or modality.

Burak Satar, Zhixin Ma, Yu-Tong Cheng et al. · 0 citations
Open access Aug 2026

Integrating Computer Vision and Large Language Models to Revitalize Dakon as Intangible Cultural Heritage

Abstract Traditional board games are increasingly digitized, yet many digital adaptations reduce embodied interaction and weaken cultural meaning. This study proposes a multimodal framework to revitalize Dakon as an intangible cultural heritage practice by integrating gesture-based computer vision and large language models. The system combines MediaPipe hand landmark detection, a Support Vector Machine classifier for gesture recognition, a deterministic Dakon game engine, and event-driven generative dialogue using Gemini 3 Flash. Experimental results show perfect precision across active gesture classes, with recall values ranging from 0.72 to 0.77, indicating that performance limitations primarily stem from hand detection robustness rather than classification accuracy. User experience evaluation involving twenty participants demonstrates high levels of engagement, perceived authenticity, clarity, and responsiveness. In addition, linguistically informed evaluation of 120 LLM-generated responses by expert evaluators reveals strong performance in cultural relevance, contextual accuracy, and motivational quality. The findings highlight that separating deterministic game logic from generative AI preserves gameplay authenticity while enabling contextual cultural interpretation. This study contributes a heritage-centered design principle in which generative AI functions as an interpretative layer rather than a rule-modifying agent. The proposed approach extends digital cultural preservation beyond static digitization toward interactive, embodied, and participatory experiences.

S. Suyahman, Cecep Hilman · 0 citations
Preprint Aug 2026

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

NARU, a benchmark designed to evaluate Narrative evolution and Reasoning on cultural Understanding in Japanese long-form video, is introduced, a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, then generates questions via task-oriented synthesis and iterative shortcut removal.

Yuheng Huang, Jianlang Chen, Jiayang Song et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.