Skip to content
Preprint

CANVAS: Consistency-Aware Navigation via Visual Adaptive Sampling for Long-Context Text-to-SVG Generation

Aug 2026 · 0 citations · 29 references
Computer Science

TL;DR

CANVAS (Consistency-Aware Navigation via Visual Adaptive Sampling), a training-free, render-aware inference framework that combines power-sharpened trajectory likelihood with visual feedback from rendered futures and derives a stroke-wise navigation rule, is introduced.

Abstract

Autoregressive large models have recently advanced Text-to-SVG generation from simple icons to complex, long-context graphics, yet standard autoregressive decoding often fails to maintain global consistency across geometry, layout, occlusion, and composition. We introduce CANVAS (Consistency-Aware Navigation via Visual Adaptive Sampling), a training-free, render-aware inference framework that combines power-sharpened trajectory likelihood with visual feedback from rendered futures and derives a stroke-wise navigation rule. It effectively estimates each candidate stroke's future value under a limited generation and rendering budget and adaptively allocates samples according to candidate uncertainty, decision influence, and rollout cost. Experiments across multiple autoregressive SVG backbones and complementary benchmarks demonstrate improvements in global consistency, which includes sound geometric relationships, spatial layouts, occlusion ordering, and overall composition, without additional training, demonstrating the effectiveness and generalization ability of our framework.

View source

Similar papers

#artificial intelligence Preprint Oct 2026

Contextual Flow Matching: Adaptive Step Selection in Flow Models for Efficient Visual Generation

Flow Matching enables high-quality visual generation via continuous-time dynamics, but inference remains costly due to multiple sequential function evaluations. Existing acceleration methods reduce the number of function evaluations but often introduce additional training overhead, degrade quality, or fail to account f...

D. J. Bajpai, Arun Verma, M. Hanawal · 0 citations
#artificial intelligence Preprint Oct 2026

SPHERE: Adaptive VR Indoor Scene Generation via LLM-Enhanced Spatial Preference Learning and Human-in-the-Loop RL

While Large Language Models (LLMs) advance 3D indoor scene synthesis, current pipelines fail to retain user-specific preferences across sessions, making immersive authoring a repetitive and physically fatiguing process. We present SPHERE, an adaptive VR generation framework that transforms isolated synthesis into conti...

Hyeonmin Lee, Zheng Wei, Kyungmin Kwon et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Weave Forcing: Compositional Memory Routing for Interactive Long Video Generation

Recent advances in autoregressive video generation have improved temporal consistency over extended durations, yet interactive storytelling requires more than continuous scene extension: a new shot may combine characters and backgrounds from different historical shots. Whole prompt retrieval can overlook the distinct r...

Zi-Yi Wang, Jun-Chi Yao, Heqian Qiu et al. · 0 citations
Preprint Sep 2026

Planning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generation

Generating text-rich images from prompts requires both textual fidelity and the coherent integration of text into the surrounding image. An explicit layout can provide structured guidance about what text should appear and where, but a well-formed plan alone does not guarantee that the renderer will realize it faithfull...

Guan-Qiao Chen, Jing Tan, Dong-Xing Mao et al. · 0 citations
Preprint Sep 2026

WorldAttention: An Efficient Attention Architecture for Interactive Video World Models

Leveraging the paradigm of autoregressive diffusion, text-conditioned interactive video world models aim to simulate temporally coherent environments guided by textual instructions. While enabling low-latency, long-duration generation is pivotal for embodied AI and simulation-based planning, current frameworks primaril...

Ze-Yu Zhang, Jin-Yuan Mao, Da-Kai An et al. · 0 citations
Preprint Aug 2026

VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation

VISTA is the first gradient-based test-time alignment framework for next-scale autoregressive image generation, and introduces the mechanisms needed to make such optimization stable across scales, together with an extensible objective space that any differentiable constraint on cross-attention can plug into.

Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.