Cloud-edge collaborative computing enables resource-constrained robots to leverage powerful cloudhosted models while retaining real-time on-device perception and control. Vision-Language Navigation in Continuous Environments (VLN-CE), which requires a robot to follow natural-language instructions through complex scenes, is a representative task that benefits from this paradigm: it relies on large Visual Language Models (VLMs) for multimodal reasoning yet demands responsive execution at the edge. However, existing VLM-based approaches remain constrained by limited context windows and insufficient planning capabilities for long-horizon tasks. We present EntityNav, an entity-centric stepwise planning framework for VLN-CE designed for cloudedge deployment. EntityNav comprises two integrated modules executed on the cloud: (1) Entity-Guided Stepwise Language Planning, which decomposes instructions into sequential, entity-centered sub-goals for explicit progress tracking, and (2) Entity-Aware Chain-of-Thought Reasoning, which generates a multi-stage structured reasoning chain whose hidden-state representations directly condition the action prediction head, regularized by a reasoning-action consistency loss. On the robot side, an edge-level module performs real-time visual capture and local trajectory refinement, with asynchronous communication overlapping cloud inference and physical motion to preserve responsiveness; an edge-side fallback mechanism further maintains safe navigation during transient cloud delays. Experiments on R2R-CE and RxR-CE benchmarks show that EntityNav achieves success rates of 62.7% and 60.3% respectively, demonstrating competitive performance against baselines. Real-world deployment on a quadruped robot further shows the framework's effectiveness under practical cloud-edge conditions.
Heng-Yi Yang, Yong Zhou, Shang Liu et al.· Fall Joint Computer Conferen...· 0 citations
Experiments show that LAUA substantially mitigates performance degradation under modality missingness across retrieval and regression tasks, attaining up to 20% relative improvement in MRR for retrieval and up to 24.3% relative improvement in MSE for regression.
Yi Wei, Xiaokai Zhou, Shanshan Feng et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.