Model the Edit, Not the Image: Visual Autoregressive Editing from a Source-Centric Perspective
This work takes a source-centric perspective on VAR editing, in which the encoded source image tokens serve as the primary visual state and the editing process focuses on condition-induced changes, and proposes EditMod, which compares source- and target-conditioned predictions under a shared autoregressive context.