Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps
Vision-Language-Motion Maps is introduced, an open-vocabulary, language-queryable 3D map queried through a rule-based intent router over open-vocabulary object nouns, in which each element carries a fused motion attribute: a VLM/LLM semantic movability prior combined with geometrically observed cross-frame motion, together with a per-element uncertainty.