This work introduces RiverVLN, to its knowledge the first benchmark designed for long-horizon USV VLN under continuous riverine motion, and PGT-NAV, a phase-grounded temporal navigation framework for USVs.
Jie-Ling Wu, Yue-Hao Huang, Jia-Jun Lv et al.· 0 citations
WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN, is presented, showing that WNM-3D outperforms strong VLM-based navigation policies and its 2D-conditioned counterpart in closed-loop navigation.
Yue-Hao Huang, Yunzi Wu, Xiaotao Zhang et al.· 1 citation
TS-Mask VLA is built upon two key designs: a Discrete Diffusion Action Expert equipped with a Bridge Attention conditioning bridge, which enables multi-layer conditioning from the VLM and facilitates more accurate and stable action generation; and a temporal-spatial 2D masking strategy for discrete action tokens that s...