To address the inherent limitations of Vision-Language Models in long-tail object retrieval for autonomous driving, this paper proposes a Dual-Granularity Structured Scene Retrieval (DG-SSR) architecture. By decoupling text queries and visual features, we introduce a parameter-free mechanism that fuses local semantic scores with macro global context. An Adaptive Negative Injection (ANI) strategy and a Soft-NCE loss further enforce fine-grained alignment and mitigate color bias. Evaluated on a curated nuScenes dataset comprising 8,500 homogeneous street views and 1,128 combinatorial queries, our method achieves an mP@5 of 33.06% with a single-query latency of 0.812 ms, outperforming CLIP (16.48%) and BLIP (23.88%) by significant margins. Extensive ablation analysis demonstrates that optimal retrieval in complex scenes is achieved through a local-dominated architecture supplemented by minimal global context.
Nan Jiang, Tongxuan Xu, Yujin Wang et al.· 2026 IEEE International Conf...· 0 citations
Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity and degrade the underlying VLMs'general-purpose capabilities after spatial fine-tuning. To this end, we propose a parameter-efficient \textit{\textbf{Spatio}-vision \textbf{L}anguage \textbf{M}odels (SpatioLM)}, that enhances spatial intelligence without extra 3D prior inputs or third-party spatial encoders. Concretely, we design a plug-and-play and non-invasive spatio-vision module that elicits the spatial knowledge inherent in VLMs. Furthermore, we innovatively leverage pseudo depth and camera information as supervision to guide the model in learning physically coherent representations. Extensive experiments show that SpatioLM achieves significant improvements in diverse tasks, including spatial perception and understanding while effectively limiting the degradation of general capabilities. Notably, the model achieves an impressive score of 71.6 on the VSI-Bench (the first model to surpass 70). In addition, it attains competitive performance when transferred to embodied manipulation tasks. Code is available at \href{https://github.com/xiaomi-research/spatio-lm}{\faGithub~spatio-lm}.
Jing Wu, Jianhua Wu, Jiayi Guan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.