ZIVIL: Zero-Shot Incremental Vision–Language Maps and Spatial Graph Representation of Construction Sites
TL;DR
This framework combines simultaneous localization and mapping (SLAM), visual-language feature extraction, incremental semantic and instance label fusion, and spatial graph construction to enable a construction robot navigation framework that supports open-vocabulary language queries.
Abstract
Construction robots are increasingly capable of performing complex, labor-intensive tasks such as bricklaying, drilling, and autonomous material handling. Using real-time perception and environmental mapping, intelligent systems can operate effectively in unstructured and dynamic site conditions that traditionally demand human expertise. Recent progress in large language models and vision foundation models offers substantial opportunities to strengthen and extend the capability of creating high-level navigational maps for construction robots. Leveraging these advances, we introduce the zero-shot incremental vision–language maps framework, which is a three-dimensional (3D) modeling system that aims to generate semantically rich map representations of construction sites in a zero-shot manner. Our framework combines simultaneous localization and mapping (SLAM), visual-language feature extraction, incremental semantic and instance label fusion, and spatial graph construction to enable a construction robot navigation framework that supports open-vocabulary language queries. Evaluation is performed on the public ConSLAM dataset, and results show that the proposed framework is capable of building a rich 3D map of columns, signs, framework, and barriers in a construction environment.