WARP-VLA is proposed, a camera-view robust VLA for diverse wrist camera configurations that adopts a Mixture-of-Experts (MoE) architecture where individual experts learn view-specific feature transformations, and a router combines them based on implicit view information.
Junmyeong Lee, Dong-Min Shin, Min-Gyu Park et al.· 0 citations
Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skelet...
Sangeyl Lee, Seunghyun Shin, S. Park et al.· 0 citations
DyMoS (Dynamic Motion Slider), a training-free and model-agnostic method that rebalances the attention pathway from generated frames to the reference frame during initial denoising steps, is proposed and demonstrated to improve motion dynamics while maintaining visual quality and fidelity to the reference image.
Wooseok Jeon, S. Park, Seunghyun Shin et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.