DynamicWAM introduces history-flow conditioning, encoding temporally aligned optical-flow frames alongside the current observation through a frozen pretrained video VAE to preserve spatial motion structure, while injecting kinematic descriptors of displacement, duration, velocity, and acceleration into the action expert to provide motion magnitude and timing.
Abstract
Dynamic manipulation requires robots to infer target motion and respond promptly, yet existing World-Action Models (WAMs) typically condition only on the current frame and execute large backbones synchronously, limiting motion awareness and responsive control in dynamic scenes. We propose DynamicWAM, a compact WAM for dynamic object manipulation with dual-path motion conditioning. DynamicWAM introduces history-flow conditioning, encoding temporally aligned optical-flow frames alongside the current observation through a frozen pretrained video VAE to preserve spatial motion structure, while injecting kinematic descriptors of displacement, duration, velocity, and acceleration into the action expert to provide motion magnitude and timing. The two complementary paths are fused through joint world-action attention. A distilled compact backbone and real-time chunking (RTC)-based asynchronous execution further enable responsive control. On DOMINO, DynamicWAM achieves a 38.2% success rate and a 53.2 manipulation score, outperforming all evaluated baselines. Across 12 real-world tasks spanning linear, circular, and compound target motion, it achieves a 46.7% average success rate, exceeding the strongest baseline by 22.9 percentage points.
MobileWAM surpasses state-of-the-art mobile manipulation policies on ManiSkill-HAB and fine-tunes to a real ARX Lift2 mobile manipulator across diverse tasks with strong generalization.
Ze-Hua Fan, Junjie He, Wenxuan Song et al.· 0 citations
DECOWAM is introduced, a whole-body world-action model that separates camera ego-motion from base and arm actions through dedicated conditional interfaces and shows that embodiment-aware factorization can support parameter-efficient joint visual prediction and whole-body control under moving viewpoints.
Siyuan Ma, Boshi Zhang, Yutian Zhang et al.· 1 citation
This work proposes Contact Flow, an embodiment-agnostic action representation that encodes manipulation through the trajectory of 3D contact points between an actor and a target object, and integrates this model into a propose-imagine-verify-act pipeline, where generated rollouts are assessed by a vision-language model before execution.
Sami Azirar, Enrico Pallotta, Jan Nogga et al.· arXiv.org· 1 citation
DreamTrajectory is presented, a trajectory-guided framework for language-conditioned mobile manipulation that introduces one component for each limitation of existing Vision-Language-Action policies, and jointly predicts an intention-level end-effector trajectory and a whole-body action chunk in a single action expert.
Zheng Yang, Wenjie Zhang, Xiang-Yu Chen et al.· 0 citations
Quadruped mobile manipulation requires two tightly coupled capabilities: reaching manipulation-ready configurations and maintaining stable contact throughout articulated-object interaction. However, existing methods often terminate navigation near the target, leaving a gap between reachability and manipulation readiness, while tracking lag, motion jitter, and contact instability limit continuous interaction. To address these challenges, we present TONAV, a unified framework integrating task-oriented navigation with action-velocity chunk learning. First, we introduce a position-velocity-coupled teleoperation framework that explicitly captures motion dynamics to improve master-follower consistency and collect smooth, temporally consistent demonstrations. Next, task-oriented navigation leverages vision-language reasoning to decompose high-level instructions into executable subgoals and adaptively refine the robot base toward a manipulation-ready configuration. Finally, action-velocity chunk learning jointly models joint positions and their temporal transitions under velocity supervision, enabling smooth and stable sustained-contact manipulation. Real-world experiments across diverse articulated-object tasks demonstrate that TONAV achieves higher success rates in both task-oriented navigation and complete mobile manipulation, mitigating the navigation-manipulation gap and improving continuous-contact interaction. The project page is at https://haochen611.github.io/TONAV.
Haoran Lin, Ming Yang, Pengfei Qi et al.· 0 citations
Experiments across action-conditioned robotic manipulation, visual planning, and model-based reinforcement learning, together with action-free driving video prediction, show that SAMPO++ improves visual prediction quality while providing stronger action alignment, counterfactual accuracy, no-op residual suppression, and long-horizon rollout consistency over strong discrete and continuous baselines.
Sen Wang, Sanpin Zhou, Huaiyi Dong et al.· IEEE Transactions on Pattern...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.