Skip to content
Preprint

Edit-VAR: Taming Visual Autoregressive Model for Precise Video Editing

Sep 2026 · 0 citations · 52 references
Computer Science

Abstract

Text-guided video editing modifies target content while preserving the appearance and temporal coherence of unedited regions. Training-based approaches provide strong control but demand substantial data and computation. Training-free methods fall into inversion-free and inversion-based paradigms. Inversion-free approaches avoid trajectory recovery, but their source-preserving guidance can limit editing strength and leave semantic changes incomplete. Inversion-based approaches recover a latent trajectory before regeneration, where approximation errors can accumulate and cause source-content drift and temporal inconsistency. We introduce Edit-VAR, the first training-free and inversion-free framework for text-guided video editing with a pretrained visual autoregressive video model. Edit-VAR directly encodes the source video into multi-scale discrete tokens and performs probability-guided conditional token replacement for source preservation. Attention-guided token-wise and scale-aware modulation selectively relaxes source constraints over edit-relevant positions and generation stages. Scale-Decoupled Generation, implemented as late-scale constraint release, regenerates motion-consistent details and reduces texture fragmentation. Residual-guided token pruning further exploits redundancy at the final two high-resolution scales to reduce inference cost. Extensive experiments and a blind user study demonstrate that Edit-VAR outperforms existing training-free video editing methods overall in editing fidelity, source preservation, temporal coherence, and inference efficiency.

View source

Similar papers

Preprint Aug 2026

Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective

This work takes a source-centric perspective on VAR editing, in which the encoded source image tokens serve as the primary visual state and the editing process focuses on condition-induced changes, and proposes EditMod, which compares source- and target-conditioned predictions under a shared autoregressive context.

Hongyi Fang, Chu-Wen Xie, Ben-Jia Zhou et al. · 0 citations
Preprint Aug 2026

EditStream: A Unified Autoregressive Framework for Interactive Video Generation and Editing

Interactive video generation and editing are becoming increasingly important for creative design. In this report, we introduce EditStream: a unified framework for interactive video generation and editing. EditStream unifies multiple video creation and manipulation tasks within a single DiT-based model through flexible...

Yu-Qian Zhou, Zhenghong Zhou, Zongze Wu et al. · 1 citation
#machine learning Preprint Sep 2026

Constrained Edit Fields for Training-Free Flow Editing

Constrained Edit Fields (CEF) achieves state-of-the-art Structure Distance, background LPIPS, and background MSE with both Stable Diffusion 3.5 Medium and FLUX, while retaining competitive instruction alignment.

Jing-Xuan Kang, Yin-Song Wang, Che Liu et al. · 0 citations
Preprint Sep 2026

Refinement Is Inherently Editable: Training-Free Prompt-to-Prompt Image Editing with Generative Refinement Network

Text-guided image editing must introduce the requested changes while preserving unrelated source content. In training-free editing, diffusion editors often use spatial controls whose inaccuracies can leave edits incomplete or alter unrelated regions. Causal autoregressive editors face a further constraint: their fixed...

Yu-Long Chen, Zi-Qian Zhang, Hao-Yu Zhang et al. · 0 citations
Conference Aug 2026

MaskFlow: Attention-Guided Localized Editing for Rectified Flow Models

Text-guided image editing using rectified flow models such as Multimodal Diffusion Transformer (DiT) has demonstrated impressive generation quality. However, existing methods apply edits globally, inevitably modifying background regions unrelated to the intended semantic change. Our key observation is that the Text-to-...

Trong-Tai Dam Vu, Vinh-Tiep Nguyen · 0 citations
Preprint Aug 2026

EditaLive! Unified Character Video Editing for Live Streaming

Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically d...

Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.