Skip to content

Vision-Language Model-Based Demonstrators for Imitation Learning in Construction Robotic Timber Assembly

Sep 2026 · Journal of Computing in Civil Engineering · Vol 41 · 64 references
Innovations in Concrete and Construction Materials

Abstract

Abstract Intelligent construction robots are deemed the future of on-site construction for improved productivity and safety. To automate the construction process, imitation learning (IL) has been adopted to train construction robots in a repertoire of tasks. However, collecting demonstrations for robots to imitate from usually requires teleoperation setup and devices, such as virtual reality (VR) and gloves. To autonomously and efficiently generate demonstration data for imitation learning in construction tasks, we propose a large foundation model-based demonstration-generation framework, in which the Segment Anything Model (SAM) is used to extract geometric representations of objects of interest and a vision-language model (VLM), conditioned on mark-based visual prompting and languages, produces high-level action suggestions. As the demonstrator framework is imperfect and is computation-intensive to be fine-tuned, only successful episodes are retained automatically and structured into demonstrations for distilling a lightweight text-conditioned robot policy via behavioral cloning (BC). We evaluate the framework on a UR10 robot arm in Isaac Sim for a language-conditioned timber assembly task with frame pose randomization. The autonomous demonstrator achieves success rates ranging from 76.9% (small perturbations) to 33.7% (largest perturbations) but can be rolled out to curate balanced demonstration data sets. Policies distilled from these demonstrations attain 100% success across all three perturbation ranges and two placement scenarios, while retaining up to 70% success under modest out-of-distribution (OOD) frame poses. An ablation on the number of demonstrations shows that the distilled policies reach 100% success with only 70% of the data in low-variation settings but require about 90% of the data to achieve 100% success under larger workspace perturbations.

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Conference Open access Dec 2013

Affordable and Energy-Efficient Cloud Computing Clusters: The Bolzano Raspberry Pi Cloud Cluster Experiment

The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.

P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al. · 110 citations · ⚡7

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.