RECAST is a robot navigation framework that combines the reasoning of a VLM with the spatial grounding of vision foundation models to build an Actionable Cost map and improves success over the strongest prior method.
Abstract
Safe and robust robot navigation across diverse environments requires a high-level understanding of complex scenes and the ability to carry it into stable motion. Recent works tackle this with learning-based models trained at scale and with approaches built on vision-language models (VLMs). However, learning-based models break down outside their training distribution, while VLM-based approaches bring that understanding but rarely ground it in the scene or align the action with it. To address these limitations, we present RECAST, a robot navigation framework that combines the reasoning of a VLM with the spatial grounding of vision foundation models to build an Actionable Cost map. Given the robot's front view and the user's instruction, we first decompose the scene with the VLM, judging which surfaces are traversable, which objects pose a risk, which heading to prefer, and which gaps are passable. Vision foundation models then ground these surfaces and objects in the image, and all four judgments are spatially recast into one compact cost map. This map both conditions the trajectory decoders and scores their proposals to select the one to execute. As the VLM's answers trail the live scene, both steps draw on cost maps from two points in time: the pivot frame the VLM judged, which carries all four judgments, and the current frame, whose terrain and collision costs are rebuilt from the current image. RECAST improves success over the strongest prior method by 13.3 points in simulation and 31.4 points on a real quadruped, and reduces the collision rate relative to it by 9.6 and 14.3 points, reaching the lowest collision rate among all methods. The project page is available at https://recast-nav.github.io/
Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be le...
Ge-Hao Zhang, Weikai Huang, S. Shailesh et al.· 0 citations
Vision-Language Navigation (VLN) in unseen indoor environments is useful in real-world robotics, where an agent must follow natural-language instructions, locate objects, and answer spatial questions without a pre-built map or fixed object vocabulary. Multimodal vision-language models (VLMs) provide strong open-vocabul...
Long Giang Vu, Cheng-Kai Yao, Yu-Xin Liu et al.· 0 citations
Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generati...
Qi-Ze Yu, Lian-Rui Fan, Bo-Yu Chen et al.· 0 citations
NaViRRator, a framework that translates user-specified start and goal locations on such maps into navigation instructions for a pretrained vision-and-language navigation (VLN) policy, is presented, supporting route-grounded language as a modular interface between human-readable maps and pretrained navigation policies.
The real-robot benchmark demonstrates that StellaVLA can use both human/robot demos and human-to-robot (XR) demos as in-context structured demonstration to help VLA model adapt to OOD tasks.
Siyu Xu, Yun-Ke Wang, Zi-Jian Wang et al.· 3 citations
This work proposes a parameter-efficient approach to fine-tune a pretrained VLM for autonomous navigation using an Imperative Learning paradigm, and introduces a unified end-to-end navigation pipeline for natural-language-driven robotic control.
Sebastian Berger, Katharina Winter, Fabian B. Flohr· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJul 30, 2026
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.