Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermor...
Chen-Zhi Liu, Yue Zhang, Jie-Hong Lin et al.· 0 citations
ART is a tool-injection framework that tunes any VLA model to leverage off-the-shelf tool modules for low-level vision, high-level affordance, and embodiment enhancement, and ART reduces the complexity of the action solution space through tool-use, which improves generalizability across different tasks but also reduces...
Ding-Ge Yi, Yanzhao Yu, Xi-Li Dai et al.· 1 citation
Long-term physical coexistence with intelligent robots requires more than capable robot policies. A persistent robotic assistant must support diverse user-facing interfaces, maintain long-horizon memory of people and preferences, coordinate across robot embodiments, and translate human intent into safe physical executi...
Weiqi Jin, Peijun Tang, Kuncheng Luo et al.· 0 citations
Deploying robot policies in the physical world requires satisfying two fundamental desiderata: reliability and smooth real-time execution. However, deploying state-of-the-art generalist models presents challenges on both fronts. Achieving the precision and robustness required for real-world deployment necessitates samp...
Guang Gao, Yuxuan Nong, Baifu Huang et al.· 1 citation
Lumo-2 is introduced, a latent world-action model that generates actions by reasoning over world dynamics in latent space that consistently outperforms strong vision-language-action and world-action model baselines, with gains on challenging real-world tasks requiring temporal reasoning, physical understanding, or high...