Skip to content

RECAST: Recasting Vision-Language Semantics into an Actionable Cost Map for Robot Navigation

Sep 2026 · 0 citations · 45 references
Computer Science

TL;DR

RECAST is a robot navigation framework that combines the reasoning of a VLM with the spatial grounding of vision foundation models to build an Actionable Cost map and improves success over the strongest prior method.

Abstract

Safe and robust robot navigation across diverse environments requires a high-level understanding of complex scenes and the ability to carry it into stable motion. Recent works tackle this with learning-based models trained at scale and with approaches built on vision-language models (VLMs). However, learning-based models break down outside their training distribution, while VLM-based approaches bring that understanding but rarely ground it in the scene or align the action with it. To address these limitations, we present RECAST, a robot navigation framework that combines the reasoning of a VLM with the spatial grounding of vision foundation models to build an Actionable Cost map. Given the robot's front view and the user's instruction, we first decompose the scene with the VLM, judging which surfaces are traversable, which objects pose a risk, which heading to prefer, and which gaps are passable. Vision foundation models then ground these surfaces and objects in the image, and all four judgments are spatially recast into one compact cost map. This map both conditions the trajectory decoders and scores their proposals to select the one to execute. As the VLM's answers trail the live scene, both steps draw on cost maps from two points in time: the pivot frame the VLM judged, which carries all four judgments, and the current frame, whose terrain and collision costs are rebuilt from the current image. RECAST improves success over the strongest prior method by 13.3 points in simulation and 31.4 points on a real quadruped, and reduces the collision rate relative to it by 9.6 and 14.3 points, reaching the lowest collision rate among all methods. The project page is available at https://recast-nav.github.io/

View source

Similar papers

Preprint Sep 2026

Grounded Action Model: 3D Grounding as a Foundation for Robotics

Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be le...

Ge-Hao Zhang, Weikai Huang, S. Shailesh et al. · 0 citations
Preprint Sep 2026

AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation

Vision-Language Navigation (VLN) in unseen indoor environments is useful in real-world robotics, where an agent must follow natural-language instructions, locate objects, and answer spatial questions without a pre-built map or fixed object vocabulary. Multimodal vision-language models (VLMs) provide strong open-vocabul...

Long Giang Vu, Cheng-Kai Yao, Yu-Xin Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generati...

Qi-Ze Yu, Lian-Rui Fan, Bo-Yu Chen et al. · 0 citations
Preprint Sep 2026

NaViRrator: Robot Navigation from Human-Readable Maps through a Learned Visual Route

NaViRRator, a framework that translates user-specified start and goal locations on such maps into navigation instructions for a pretrained vision-and-language navigation (VLN) policy, is presented, supporting route-grounded language as a modular interface between human-readable maps and pretrained navigation policies.

Ayun Lee, Jiseon Kim, Giseop Kim · 0 citations
Preprint Sep 2026

"Dear LLaVA, Please Drive": A Depth-Aware Vision-Language Agent for Closed-Loop Robotic Control

This work proposes a parameter-efficient approach to fine-tune a pretrained VLM for autonomous navigation using an Imperative Learning paradigm, and introduces a unified end-to-end navigation pipeline for natural-language-driven robotic control.

Sebastian Berger, Katharina Winter, Fabian B. Flohr · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.