Skip to content

AnnoSketch: Evaluating and Collecting Human Sketches for MLLM-assisted Chart Annotation

Sep 2026 · 0 citations · 56 references
Computer Science

TL;DR

This study examines when sketch input is useful for MLLM-generated chart annotations across variation in chart type and caption type, and presents AnnoSketch, a database of annotation sketches collected across 160 chart-caption pairs from the conditions in which sketch guidance proved most beneficial.

Abstract

As multimodal large language models (MLLMs) support a growing range of input modalities, increasing work explores how to incorporate rough sketches to convey user intent. For annotated chart generation, it remains unclear what annotation sketches people provide and when such visual input helps MLLMs generate more useful annotations. In this study, we examine when sketch input is useful for MLLM-generated chart annotations across variation in chart type and caption type. In addition, we qualitatively analyze participants'explanations of their output preferences to characterize what made generated annotations more or less helpful. To further document participants'annotation sketches, we present AnnoSketch, comprising 1,600 annotation sketches collected across 160 chart-caption pairs from the conditions in which sketch guidance proved most beneficial, together with participants'annotation intents, perceived comprehension difficulty, and self-reported expressive limitations. We also label these sketches with structured metadata describing how each sketch relates to its caption and how participants express annotations through visual marks. Together, our study and AnnoSketch help determine when to solicit sketch input and provide empirical source for how people sketch chart annotations to support captions. The dataset and supplemental materials are available in our OSF repository.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al. · 1 citation
#human-computer interacti... Preprint Sep 2026

From Images to Tasks: Characterizing Multimodal LLM Interactions in the Wild

Multimodal large language models (LLMs) increasingly integrate vision and text, yet how people use them in natural settings remains underexplored. We seek to answer the question: when users upload images, what tasks are they trying to accomplish? Analyzing over 40,000 de-identified image-upload conversations from Micro...

Jin-Yi Ye, Scott Counts, Gaurav Verma et al. · 0 citations
#natural language process... Preprint Sep 2026

CALICO: A Human-Centered, Codebook-Aligned System for Annotation

CALICO is presented, a human-centered, codebook-aligned annotation workflow that treats prompts as editable, versioned, and optimizable artifacts and integrates codebook parsing, prompt generation, result inspection, prompt versioning, natural language human feedback, and label-supervised prompt optimization through ex...

Boqin Yuan, Xiao-Yi Gu, F. Li et al. · 0 citations
Preprint Aug 2026

Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization

This work presents AudioChaps, a post-training framework for aligning end-to-end LALMs for this task via Group Relative Policy Optimization (GRPO) guided by Chain-of-Thought (CoT) reasoning, and demonstrates that GRPO-trained LALMs can reliably transform unstructured auditory streams into navigable, structured media.

Tony Alex, Wish Suharitdamrong, Sara Atito et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.