World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action...
Cheng-Tao Lv, Jin-Yang Du, Shu-Yi Feng et al.· 0 citations
The proposed MaskFlow, a training framework for precise localization, consistent background preservation, and seamless boundary transitions, incorporates the mask into the probability path and flow-matching objective, coordinating generation within the editable region with source preservation outside it.
Rui Xu, Yang Yong, Shun-Zi Yang et al.· 0 citations
The experiments show that constrained decoding consistently enforces syntactic validity, but does not reliably improve semantic accuracy and may even degrade performance for smaller models or complex grammars, and reveal clear task-and model-dependent boundaries for effective constrained decoding.
Xiao-Kun Xiong, Zhengjie Xu, Junyi Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.