Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-re...
Guo-Qing Ma, Ming-Qi Yuan, Chen Gao et al.· 0 citations
FAN is introduced, which achieves the highest performance and demonstrates consistent robustness, providing insightful guidance for building stable action representations in achieving effective lifelong VLA adaptation.
Yi-Jun Hong, Jia-Run Zhu, Xiao-Quan Sun et al.· 0 citations
This work proposes UCA-Flow, a unified condition-action modeling framework for accurate one-step action generation that unifies observation conditions, timestep conditions, interval conditions, and action tokens into a single sequence, and processes them with a Unified Condition-Action Transformer for joint condition-a...
Xin-Yu Zhou, Zikun Cai, Kuangji Zuo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.