Building on recent advances in world models, World Action Models (WAMs) jointly model video prediction and action generation. However, they typically represent videos in 2D pixel space, creating a representation gap with 3D space in which robotic actions are executed. Recent 3D approaches introduce 3D information, but fail to fully exploit the dynamics of 3D structures. In this work, we propose 4D-WAM, a model-agnostic training strategy that injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment. To this end, we introduce two complementary objectives: 1) motion alignment, which aligns temporal feature variations across adjacent frames and encourages the model to build local 4D awareness during training, and 2) destination alignment, which guides the model to infer the final destination from the source frame by minimizing the gap between their attention-like similarity distributions. Together, these objectives provide both local motion supervision and long-horizon goal guidance, enabling WAMs to learn trajectory-level spatiotemporal representations. Extensive in-distribution and out-of-distribution experiments across different base models demonstrate the model's improvements in spatial understanding, execution precision, robustness, generalization, and versatility.
Lishan Yang, Wenxuan Song, Xi Wang et al.· 1 citation
Autonomous exploration in large-scale and cluttered environments remains a severe challenge for uncrewed aerial vehicles (UAVs). Existing methods still suffer from high computational costs, long-distance revisits, and discontinuous flight. To address these issues, we propose PACE, a hierarchical exploration framework designed for high-speed flight in large-scale and cluttered environments. First, we introduce a passability-based frontier classification mechanism to distinguish potential channels from local dead-ends, providing essential priors for decision-making. On this basis, a memory-guided global planner is proposed, which maintains the consistency of long-term global intent with low computation cost through an adaptive sliding window and anchor point constraints. Furthermore, it employs an adaptive priority adjustment method to adjust the global order more reasonably in scenarios of various scales to avoid future revisits. Finally, an intent-aware local planner is proposed to achieve agile and fluid flight by switching between traverse and link modes, tightly integrating global guidance with local maneuvers. Extensive simulation experiments demonstrate that the proposed method significantly outperforms SOTA methods in flight velocity and exploration efficiency. Real-world experiments further validate the value of the proposed method in practical applications.
Yuxiang Gao, Zhuoxuan Wang, Xianlu Tao et al.· IEEE Robotics and Automation...· 0 citations
PSG-JEPA is proposed, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes.
Haodong Yan, Jiaguang Zhu, Ming-Ming Jia et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.