Skip to content
Preprint

Language-Grounded Semantic Target Navigation for Autonomous Surface Vehicles

Sep 2026 · 0 citations · 75 references
Computer Science

TL;DR

Results indicate that the proposed perception-to-control framework can support language-grounded target approach manoeuvres of ASV under the complex port environments and demonstrate the importance of semantic grounding, harbour-aware filtering, and semantic verification for reliable language-grounded ASV navigation.

Abstract

Autonomous Surface Vehicles (ASVs) are increasingly expected to operate in ports and harbour environments, where operators may specify navigation targets through language-based descriptions rather than predefined coordinates or fixed target identifiers. However, existing ASV navigation methods mainly execute predefined geometric goals or task-specific objectives and give limited attention to language-grounded target specification. This study proposes Semantically Grounded Navigation (SGNav), a framework that enables an ASV to identify and approach a maritime target from an operator-provided description. SGNav integrates text-guided semantic grounding, harbour-aware candidate filtering, CLIP-based semantic verification, grounded target control-state construction, and Proximal Policy Optimisation-based closed-loop control. It grounds the target description in onboard RGB observations, suppresses visually or semantically irrelevant distractors, and converts the selected target into a compact control-oriented representation for policy execution. Experiments in simulated port environments show that SGNav achieves success rates of $97.0\pm1.2\%$, $92.0\pm1.5\%$, and $90.0\pm1.8\%$ across three representative target-reaching tasks, with semantic target accuracy above $97\%$ and wrong-target rates below $3\%$. SGNav also maintains $97.7$--$98.7\%$ success rates across held-out port layouts. In the Task~3 ablation study, removing harbour-aware filtering or semantic consistency reduces the success rate to $40.4\pm2.6\%$ and $50.4\pm3.1\%$, respectively. These findings demonstrate the importance of semantic grounding, harbour-aware filtering, and semantic verification for reliable language-grounded ASV navigation. These results indicate that the proposed perception-to-control framework can support language-grounded target approach manoeuvres of ASV under the complex port environments.

View source

Similar papers

Preprint Aug 2026

From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation

This work presents an instruction-grounded semantic enhancement module that injects object-level semantics and relative spatial cues into the current observation state, and develops a relevance-aware dynamic temporal aggregation strategy that reweights the full history buffer while converting a few high-relevance frame...

Ze-Yuan Ma, Jiaxin Chen, Di Huang · 0 citations
Preprint Sep 2026

Map the Possibilities: Spatial Belief Fields for Language-Goal Aerial Navigation

Language-goal aerial navigation requires an agent to local- ize a potentially unobserved target from relational instruc- tions and partial observations, and translate this inference into metric actions in large-scale continuous environments. Existing methods often reduce language grounding to one single waypoint or act...

Hao-Tian Xu, Yue Hu, Zheng-Qiu Zhu et al. · 0 citations
Conference Aug 2026

From Language to 3D Search Goal: A Lightweight Closed-Loop Framework for Language-Guided UAV Target Search

A key problem in language-guided UAV target search is how to transform a language-referred target in the current observation into an executable spatial goal. Existing methods either predict actions directly or introduce relatively heavy mapping, memory, or planning modules, making the intermediate link between semantic...

Jun-Song Zhang, Yao-Hong Zhang, Rui Guan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

A dual-layer semantic-spatial belief mapping framework that transforms transient VLM observations into persistent spatial guidance is proposed, and object-conditioned visual reasoning with conservative evidence qualification is introduced to improve observation reliability before spatial accumulation.

Jian-Qiang Xiao, Xiang Deng, Yue-Xuan Sun et al. · 0 citations
Preprint Sep 2026

PIVOT: Physically Informed Vision-Language Off-Road Traversability for Field Robot Navigation

Terrain assessment is a critical capability for off-road mobile robots, enabling safe and reliable navigation through unstructured and geometrically complex environments. Conventional geometry-based terrain assessment is fast to compute but often overly conservative in unstructured environments. We present PIVOT: a Phy...

Ao-Ran Jiao, Wen-Da Zhao, Hshmat Sahak et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.