Skip to content
Preprint

VectorHarness: Recovering Editable, Relation-Preserving Structure from Scientific Graphics

Sep 2026 · 0 citations · 45 references
Computer Science

TL;DR

VectorHarness is presented, a multi-agent framework for raster-to-authoring reconstruction that recovers heterogeneous components using type-appropriate native representations that recovers heterogeneous components using type-appropriate native representations.

Abstract

Converting scientific graphics into editable representations remains a challenging problem for image-to-code generation because of their heterogeneous elements and complex layouts. Recent multi-agent reconstruction systems have advanced this line of work, but often follow a copy-paste paradigm: the reconstructed image closely resembles the original, while complex regions remain effectively uneditable. We instead formulate a different objective, raster-to-authoring reconstruction, which aims to recover an authoring representation that supports native, customized editing rather than mere visual replication. To this end, we present VectorHarness, a multi-agent framework for raster-to-authoring reconstruction that recovers heterogeneous components using type-appropriate native representations. Text, formulas, shapes, connectors, icons, charts, and tables are reconstructed as natively editable objects, while intrinsically image-based regions remain raster content. To systematically evaluate reconstruction quality, we introduce VectorHarness-Bench, which jointly assesses rendering fidelity, raster fallback coverage, executable object edits, and relation-preserving edits. Experiments show that VectorHarness improves executable edit success and relation preservation, reduces avoidable raster fallback, and maintains high visual fidelity across heterogeneous graphics.

View source

Similar papers

Preprint Sep 2026

Back2Struct: Making Structured Images Editable Again

Structured images, such as diagrams, charts, and flowcharts, are inherently symbolic and can be compactly represented in an editable format, yet in practice, they are often rendered as images, and therefore not graphically editable. This mismatch presents a significant challenge for researchers, engineers, and designer...

Peng-Yu Yan, Yi-Xin Wu, Yun-Jie Tian et al. · 0 citations

SVG-BERT: Code-Native Representations for Structure-Aware Vector Graphics Workflows

SVG-BERT, a code-native encoder pretrained directly on SVG source code as representation infrastructure for structure-aware vector design workflows, consistently outperforms generic text encoders and provides information complementary to raster encoders, which remain stronger on strongly visual semantic tasks.

Boris Malashenko, Ivan Jarsky, V. Shalamov et al. · 0 citations
Preprint Aug 2026

Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ

This work introduces Edit2TikZ, a comprehensive benchmark for scientific figure editing tasks, featuring 1,548 diverse and high-quality samples, and constructs a human-aligned evaluation framework to measure whether a requested edit is completed while irrelevant content is preserved.

Zong-Yun Zhang, Jiacheng Ruan, Xian Gao et al. · 0 citations
Preprint Aug 2026

AcroMELD: Recovering Interactive PDF Forms with Structure-Aware Graph Set Transformers

Interactive PDF form fields are often absent from documents that visually resemble forms, leaving users unable to enter data without printing or external editing tools. Detecting the missing widgets is difficult because a field may be indicated by several overlapping cues, born-digital PDFs expose useful but incomplete...

Samuel Abramov · 0 citations
#artificial intelligence Preprint Sep 2026

Figures as Programs: Recursive Generation of Editable Scientific Figures

Examination of the proposed FigTree system produces high-quality figures, while also enabling more effective editing than existing raster-based methods, and proposes a multi-agent system that automatically transforms a scientific paper into a structured vector figure.

Yepeng Liu, Da-Sen Dai, Cheng-Zhi Liu et al. · 2 citations
Preprint Sep 2026

What Makes a 3D Scene Editable? A Factorized Benchmark of Fidelity, Locality, Consistency, and Preservation

Neural 3D scene editing is often evaluated by semantic alignment alone, although a convincing result may alter unrelated content or become inconsistent across views. We introduce EditBench3D, a representation-agnostic benchmark that treats editing as controlled information replacement. It evaluates four complementary p...

Sariah Patro, Arjun Mehra, Nikhil Bhatia · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.