This work introduces BioVLN, a simulation platform for developing and evaluating visual-language navigation agents in biomedical laboratories and shows that geometric exploration reaches 74.4--87.5% success, while sampling multiple valid positions in the operation area improves success and reduces unsafe proximity.
Zhe Liu, Quan Lu, Zhao-Hui Du et al.· arXiv.org· 0 citations
LightNav-0 is presented, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads, and establishes compact VLMs as a unified and transferable backbone for generalist embodied navigation.
Shao-An Wang, Ao-Cheng Luo, Fei Huang et al.· 2 citations
This paper presents the first systematic benchmark of 17 edge-deployable SLMs against 4 online APIs for robotic navigation instruction decomposition, and proposes a lightweight hybrid semantic-geometric goal localization framework that combines open-vocabulary object detection, prompted segmentation, and LiDAR geometry...
Ali Salmasi, Xian-Jia Yu, Tomi Westerlund· arXiv.org· 0 citations
Visual Language Navigation (VLN) enables robots to follow natural language instructions to navigate visually perceived environments. Typically, VLN systems are trained on multi-modal datasets that pair visual scenes with navigation instructions. While prior work has focused on generalising to unseen environments, lingu...
Malak Sayour, Pamela Carreno-Medrano, Michael Burke et al.· IEEE Robotics and Automation...· 0 citations
Long-term navigation for service robots faces crit- ical challenges like the accumulation of odometry drift and sensor error, which progressively degrade 2D maps and renders traditional path planning algorithms (e.g., A*, RRT*, DiPPer, ViT-A*) ineffective over time. To address this, we propose a user-friendly, interact...
Praveen Kumar, K. Guruprasad, Tushar Sandhan· 0 citations
Experimental results in various task scenarios show that the proposed framework consistently improves overall task success rates compared with unimodal settings with different LLMs and achieves a higher success rate compared to using only visual or force data.