2026· Annual Meeting of the Association for Computational Linguistics· pp. 21858-21873· 0 citations· 30 references
Computer Science
TL;DR
The first comprehensive benchmark dedicated to deciphering diverse Ancient Chinese musical notations, including five types of ancient Chinese music notation systems, is introduced, highlighting the limitations of current MLLMs in processing diverse spatial-symbolic dependencies.
Abstract
Multimodal Large Language Models (MLLMs) excel in general tasks but struggle with specialized, structured cultural symbols. We introduce BoYaEval, the first comprehensive benchmark dedicated to deciphering diverse Ancient Chinese musical notations, including five types of ancient Chinese music notation systems. These systems utilize unique spatial layouts and specialized ideograms to encode pitch and intricate playing techniques. BoYaE-val comprises 3,175 high-quality images across these notation styles and establishes a three-tier evaluation: Structural Parsing (symbol recognition), Instructional Translation (technique mapping), and Musical Reasoning (melody derivation). We evaluate 21 leading MLLMs. Results indicate that while models perform adequately in basic recognition, they fail in cross-system compositional logic, scoring only around 27% on reasoning tasks. BoYaEval highlights the limitations of current MLLMs in processing diverse spatial-symbolic dependencies, bridging the gap between ancient wisdom and modern AI for digitizing intangible cultural heritage. The BoYaEval benchmark is publicly available at https://huggingface. co
Experimental results indicate that fully automated data curation combined with imbalance-aware training yields non-trivial improvements, but models still struggle to capture fine-grained acoustic features, indicating a gap between surface-level alignment and deep musical comprehension.
Ziya Zhou, Shangda Wu, Shenyang Xu et al.· 0 citations
Abstract Traditional board games are increasingly digitized, yet many digital adaptations reduce embodied interaction and weaken cultural meaning. This study proposes a multimodal framework to revitalize Dakon as an intangible cultural heritage practice by integrating gesture-based computer vision and large language models. The system combines MediaPipe hand landmark detection, a Support Vector Machine classifier for gesture recognition, a deterministic Dakon game engine, and event-driven generative dialogue using Gemini 3 Flash. Experimental results show perfect precision across active gesture classes, with recall values ranging from 0.72 to 0.77, indicating that performance limitations primarily stem from hand detection robustness rather than classification accuracy. User experience evaluation involving twenty participants demonstrates high levels of engagement, perceived authenticity, clarity, and responsiveness. In addition, linguistically informed evaluation of 120 LLM-generated responses by expert evaluators reveals strong performance in cultural relevance, contextual accuracy, and motivational quality. The findings highlight that separating deterministic game logic from generative AI preserves gameplay authenticity while enabling contextual cultural interpretation. This study contributes a heritage-centered design principle in which generative AI functions as an interpretative layer rather than a rule-modifying agent. The proposed approach extends digital cultural preservation beyond static digitization toward interactive, embodied, and participatory experiences.
S. Suyahman, Cecep Hilman· Preservation, Digital Techno...· 0 citations
PUMA (Polish Unified Multimodal Assessment) is proposed, a novel benchmark of 900 hand-crafted tasks designed to probe the limits of multimodal models in the Polish cultural and linguistic context and open-source the evaluation framework to advance localized multimodal AI research.
Slawomir Dadas, Michał Perełkiewicz, Rafal Poswiata et al.· 0 citations
BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents, is introduced, and existing hallucination detection methods are compared.
L. Chubarova, A. Kuleshova, D. P. Volkov et al.· 0 citations
This study introduces MVSCBench, the benchmark designed to evaluate multimodal AI models’ ability to interpret symbolic and cultural meanings in music videos. Using K-pop music video clips, we organize music video understanding into three stages: surface perception, symbolic interpretation, and cultural grounding. Based on this framework, we construct a dataset of 293 video clips and 3,516 annotations, covering space, artist, character symbolism, object symbolism, literary symbolism, social-historical symbolism, and fandom cultural symbols. The results show that current multimodal AI models perform well on surface perception tasks, but still face clear limitations in symbolic interpretation and cultural grounding. In particular, the models often struggle to correctly identify relevant visual cues and frequently produce hallucinated interpretations in categories involving deeper meanings. MVSCBench provides a new benchmark for evaluating symbolic and cultural understanding in music videos and offers a useful framework for future research.
This paper presents a rigorous comparative study of generative and encoder-based neural architectures for NER on all eleven languages of the Naamapadam benchmark; identifies three language clusters--encoder-dominant, partial-coverage, and failure-zone; and provides actionable deployment guidelines grounded in transfer learning and low-resource NLP principles.
Jakkala Mahesh, Jatavath Shravan Kumar, K. Shivani et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.