Skip to content
Preprint

ChartAnno: Benchmarking Multimodal Large Language Models for Chart Annotation Generation

Aug 2026 · 0 citations · 54 references
Computer Science

TL;DR

This work introduces ChartAnno, a comprehensive benchmark for evaluating MLLMs on chart annotation generation, and develops a multidimensional evaluation framework combining rule-based and LLM-judged metrics to assess execution, structural compliance, semantic consistency, and design effectiveness.

Abstract

Annotations are essential to communicative visualization, helping explain data, emphasize key findings, and guide attention. While multimodal large language models (MLLMs) offer new opportunities for automatic chart annotation authoring, their capabilities in this task remain underexplored. To address this gap, we introduce ChartAnno, a comprehensive benchmark for evaluating MLLMs on chart annotation generation. ChartAnno contains 1,200 real-world charts with paired annotated and unannotated executable code, along with 3,600 annotation instructions spanning three levels of specificity. We also develop a multidimensional evaluation framework combining rule-based and LLM-judged metrics to assess execution, structural compliance, semantic consistency, and design effectiveness. We evaluate 10 representative MLLMs under two primary chart input settings: (1) chart code alone and (2) both code and chart image. Results reveal that proprietary models lead overall, though open-source models narrow the gap. While higher instruction specificity improves annotation quality, inferring abstract communicative intent remains difficult across all models. Providing chart images yields marginal benefit when code is available. We also examine the effect of chart code through an image-only ablation and analyze the effects of multiple task complexity indicators and instruction-level transitions. Further analyses characterize common failure modes and validate the reliability of the LLM-based judge. Experiments with D3 and SVG demonstrate the generalizability of ChartAnno beyond its primary Python setting.

View source

Similar papers

Sep 2026

Advancing Scientific Chart Understanding: The SCI-CQA Benchmark and Beyond.

In real-world applications, current multimodal large models are often overestimated in their ability to understand scientific charts. To assess their true capabilities and identify key performance bottlenecks, we conducted an in-depth study on scientific chart understanding. Charts in scientific literature often featur...

Ling-Dong Shen, Qigqi, Kun Ding et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al. · 1 citation
#natural language process... Preprint Aug 2026

DVBench: Benchmarking MLLMs for Understanding Dynamic Charts and Narratives in Data Videos

DVBench is introduced, a benchmark for evaluating MLLMs on data videos, a storytelling medium that integrates dynamic charts with structured narratives, and decompose data video understanding into five dimensions, identifying two notable phenomena.

Bomiao Wang, Zekai Shao, Jiexiang Lan et al. · 0 citations
Preprint Sep 2026

ChartJudgeBench: Evaluating LMM Judges for Chart-to-Code Generation

Building strong chart-to-code systems increasingly relies on reinforcement learning, whose effectiveness depends critically on the quality of the reward signal. Large Multimodal Models (LMMs) play a natural critical role in jointly assessing chart visual appearance and task requirements. They are therefore increasingly...

Li-Jian Wu, Henry Hengyuan Zhao, Zi-Jian Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via Code

Chart editing requires cross-modal edit grounding, realizing a requested visual change in the code that draws it, with necessary related updates and without altering unrelated content. Existing benchmarks emphasize either code executability or chart quality, but their metrics do not clearly distinguish request completi...

Jia-Xiang Tang, Yi Zhou, Chad DeLuca et al. · 0 citations
Preprint Aug 2026

Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ

This work introduces Edit2TikZ, a comprehensive benchmark for scientific figure editing tasks, featuring 1,548 diverse and high-quality samples, and constructs a human-aligned evaluation framework to measure whether a requested edit is completed while irrelevant content is preserved.

Zong-Yun Zhang, Jiacheng Ruan, Xian Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.