2026· International Conference on Language Resources and Evaluation· pp. 11085-11107· 0 citations· 45 references
Computer Science
TL;DR
STRUDEL establishes a scalable foundation for assessing and advancing VLMs torward deeper and more systematic understanding of structured visual information across domains, and reveals that models excel at association tasks, yet struggle with quantification and identification, where precise structural understanding is required.
A framework for generating large-scale diagram-grounded instruction data by leveraging terminology derived from scientific curricula is introduced, and augmenting existing models such as LLaVA OneVision with SciGram establishes new state-of-the-art performance on diagram question answering.
SABRE is established as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark, and the results establish SABRE as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark.
Zi-Xuan Lan, Luzhe Sun, Matthew R. Walter et al.· 0 citations
AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.
Han Bao, Yue Huang, Yan-Bo Wang et al.· Proceedings of the 32nd ACM...· 0 citations
benchmark provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbolic reasoning, and provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbol...
Xinqi Yang, Kang An, Tengyue Wang et al.· 0 citations
Two novel contributions are introduced: CodeEval and CodeQual, an open-source execution framework that provides researchers with a ready-to-use evaluation pipeline for evaluating and improving LLMs in software engineering contexts, encompassing both functional correctness assessment and subjective code quality evaluati...
Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT forces models to serialize visual problems into awkward prose. Although architectural solutions exist...
Tsung-Han Wu, Heekyung Lee, An-Ya Ji et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.