Skip to content
Open access

STRUDEL: Unrolling a Benchmark for Evaluating Vision-Language Models on Structured Diagram Understanding across Domains

2026 · International Conference on Language Resources and Evaluation · pp. 11085-11107 · 0 citations · 45 references
Computer Science

TL;DR

STRUDEL establishes a scalable foundation for assessing and advancing VLMs torward deeper and more systematic understanding of structured visual information across domains, and reveals that models excel at association tasks, yet struggle with quantification and identification, where precise structural understanding is required.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

From Terminology to Diagrams: Visual-Instruction Generation for Scientific Diagram Understanding

A framework for generating large-scale diagram-grounded instruction data by leveraging terminology derived from scientific curricula is introduced, and augmenting existing models such as LLaVA OneVision with SciGram establishes new state-of-the-art performance on diagram question answering.

Raúl Ortega, José Manuél Gómez-Pérez · 0 citations
Review Aug 2026

SABRE: Scalable and Automated Benchmarking of VLMs under Stress

SABRE is established as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark, and the results establish SABRE as a reusable framework for constructing and refreshing VLM stress tests rather than a single fixed benchmark.

Zi-Xuan Lan, Luzhe Sun, Matthew R. Walter et al. · 0 citations
Book Open access Aug 2026

AutoDavis: Automatic and Dynamic Evaluation Protocol of Large Vision-Language Models on Visual Question-Answering

AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.

Han Bao, Yue Huang, Yan-Bo Wang et al. · 0 citations
Preprint Aug 2026

CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits

benchmark provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbolic reasoning, and provides a focused testbed for measuring whether multimodal models can transform technical visual evidence into sustained, physically valid symbol...

Xinqi Yang, Kang An, Tengyue Wang et al. · 0 citations
Open access

Evaluation and Distillation of Source Code Generation Tasks by Large Language Models

Two novel contributions are introduced: CodeEval and CodeQual, an open-source execution framework that provides researchers with a ready-to-use evaluation pipeline for evaluating and improving LLMs in software engineering contexts, encompassing both functional correctness assessment and subjective code quality evaluati...

Danny Brahman · 0 citations
#computer vision Preprint Aug 2026

CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions

Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT forces models to serialize visual problems into awkward prose. Although architectural solutions exist...

Tsung-Han Wu, Heekyung Lee, An-Ya Ji et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.