Skip to content

Semantic scene graphs for robot-based maintenance and inspection of civil infrastructure

· 0 citations · 27 references

TL;DR

The results demonstrate the capability of the proposed framework to support scene understanding, generate scene graphs, and enable LLM-based task planning for robot-based maintenance and inspection of civil infrastructure.

View source

Similar papers

Review Open access Jul 2026

Research on Scene Perception and Task Semantic Understanding for Indoor Service Robots: Based on Literature Review and Case Analysis

The study argues that scene perception and task semantic understanding constitute a continuous intelligent chain from "environment recognition" to "task execution" so as to enhance their practical performance and operational reliability.

Haoxiang Huang · 0 citations
Preprint Sep 2026

A hybrid pipeline for dynamic ontology-based semantic mapping

Semantic mapping plays a crucial role in the ability of a robot to interact with objects, operate and navigate a complex environment. The most common pipeline for semantic mapping consists of geometric mapping and localization (SLAM), perception, semantic fusion and semantic representation. However, more recent works also integrate a form of prior knowledge in their application, most notably knowledge graphs or semantic scene graphs, to improve contextual understanding of the environment. In this paper, we present a hybrid pipeline for semantic mapping. Our system incorporates an external calibrated camera using homography projection for geometric mapping and localization, combined with object detection, persistent object tracking and ontology driven semantic updates to build a dynamic semantic world model. Linear regression models are also used for correction of the estimated values of real world coordinates. The system continuously updates object instances, spatial properties and semantic relations based on real time sensory data. Ontologies are selected as form of knowledge representation due to their hierarchical structure, semantic expressiveness and support for dynamic world modelling.

Konstantinos Dimitropoulos, Ioannis Hatzilygeroudis · 0 citations

ZIVIL: Zero-Shot Incremental Vision-Language Maps and Spatial Graph Representation of Construction Sites

This framework combines simultaneous localization and mapping (SLAM), visual-language feature extraction, incremental semantic and instance label fusion, and spatial graph construction to enable a construction robot navigation framework that supports open-vocabulary language queries.

Charles M. Raines, I. Fernandez, Mandy Sun et al. · 0 citations
Open access 2026

BIM-GRASP: A Graph-RAG Approach for Semantic Parsing of IFC Building Models

Building Information Modeling (BIM) is central to modern construction and design, with the Industry Foundation Classes (IFC) format serving as a widely adopted open standard for representing building models. However, IFC data is notoriously complex as it encodes structural, geometric, and semantic information in deeply nested relationships that are difficult to interpret without specialized tools and domain expertise. This complexity limits access to meaningful insights, particularly for non-technical stakeholders. We introduce BIM-GRASP, a novel system-level application that combines Graph-Retrieval Augmented Generation (Graph-RAG) with Generative AI to enable natural language interaction with IFC building models. BIM-GRASP transforms IFC files into a knowledge graph, which is then queried by a Large Language Model (LLM) to answer questions about building elements such as geometric attributes, material specifications, and inter-element relationships. Unlike traditional IFC parsers, BIM-GRASP supports complex queries that require reasoning across multiple layers of the IFC hierarchy as it is able to retrieve information from disparate parts of the model. This allows users to ask questions and receive accurate, context-aware answers in natural language, without relying on schema-level navigation or technical syntax. Our findings show that providing the model with domain-specific In-Context Learning (ICL) significantly improves precision and relevance in information extraction, achieving an average accuracy of 88% across diverse building models and query types. To the best of the authors’ knowledge, BIM-GRASP is the first framework of its kind to enable the parsing of IFC data as a knowledge graph through natural language without requiring specialized technical expertise or manual schema navigation. By bridging the gap between complex building data and non-technical stakeholders, BIM-GRASP accelerates decision-making and improves transparency across the built environment.

Hadeel Saadany, S. Iranmanesh, Malik U. Mehmood et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.