Skip to content
Open access

Human–robot collaboration in building disassembly: a multi-agent LLM architecture

Jul 2026 · Construction Robotics · Vol 10 · 0 citations · 22 references

TL;DR

This proof of concept demonstrates that agentic multi-agent LLM systems can enable adaptive human–robot collaboration under uncertainty, beyond natural language interfaces through integrated domain knowledge, physics validation, and agentic reasoning.

Abstract

Building disassembly is critical for circular economy material reuse, yet remains rare due to cost and safety constraints, leading to demolition and material downcycling. Automation could improve both efficiency and safety, but currently available technology does not yet enable full automation. We propose a human–robot collaboration system architecture that uses agentic large language models. We test this approach in building disassembly—an unstructured, safety–critical domain where conventional pre-programmed robotics are inadequate. The agentic architecture combines curated domain knowledge, physics simulation for stability validation, and natural language interfaces, enabling the robot to participate through proactive reasoning rather than follow control commands. We evaluated the architecture through three progressively complex scenarios: collaborative spatial adaptation, collaborative decision-making, and learning. The main contribution is a modular, data-grounded HRC methodology in which specialized LLM agents perform agentic reasoning: the robot assesses situations, retrieves relevant procedural knowledge, validates decisions through simulation, and negotiates solutions with human operators. This proof of concept demonstrates that agentic multi-agent LLM systems can enable adaptive human–robot collaboration under uncertainty, beyond natural language interfaces through integrated domain knowledge, physics validation, and agentic reasoning.

Read PDF

Similar papers

Preprint Aug 2026

Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs

Agentic AI frameworks interpret open-ended task goals and decompose them into multi-step plans. Richer information about embodiment-specific capabilities, physical preconditions, and cross-robot coordination improves grounding, but does not eliminate infeasible, mistimed, or unsafe physical actions. Physical robot crews therefore require an explicit architectural interface between semantic planning and execution, where every planned action is verified against robot capabilities, system state, and workflow constraints before actuation. This paper introduces Physical Agentic AI, a framework for skill-grounded robot agent orchestration, in which each robot exposes a typed library of executable skills while a foundation model planner decomposes a task into phases and assigns each phase to a robot-skill pair. A Robot Orchestration layer exposes the skill library, robot state, named locations, and workflow contracts to a non-actuating Mission Planner, while a deterministic Robot Orchestrator validates and authorizes one skill at a time. We evaluate on a drone-UGV search-and-dispatch mission, where every mission in every condition is executed live in Gazebo, and on a humanoid-quadruped transportation task using hardware-equivalent skill interfaces plus two physical trials on a Unitree G1 and Go2. Varying planner knowledge and runtime enforcement independently, we find that retrieval raises skill grounding from 51% to 96% yet leaves informed planners dispatching 23-29% of faulted steps. Per-dispatch enforcement reduces false dispatch to 0% with no false blocks, and a held-plan ablation confirms that the gate, not plan variation, is responsible. Live execution makes the difference physical: without enforcement all eight injected faults crossed the orchestration boundary and six produced robot motion; with enforcement all eight were refused before motion.

Xinyuan Liu, Eren Sadikoglu, R. Chatterjee et al. · 0 citations
Preprint Sep 2026

Harness Robotic OS: A Unified Embodied-Agent Runtime for Closed-Loop Quadruped Inspection

Autonomous property inspection requires more than robust robot navigation: a deployable system must connect heterogeneous sensing, reusable autonomy capabilities, multimodal scene understanding, human interaction, and enterprise response within a traceable operational loop. Existing quadruped inspection systems commonly integrate these functions through task-specific interfaces, making contextual coordination, knowledge reuse, and controlled adaptation difficult. This paper presents \textit{Harness Robotic OS} (HROS), a unified embodied-agent runtime, and Argos, its realization for residential-community inspection. HROS organizes the system into robot runtime, embodied autonomy skills, cognitive agent runtime, and interaction and operations planes. A shared context connects physical state with agent reasoning; streaming ASR/TTS supports voice-based mission interaction; hierarchical working, episodic, and semantic memory preserves operational knowledge; and a safety-gated self-evolution loop converts execution traces into versioned candidate updates without permitting unconstrained online modification. The Argos prototype integrates a Vbot quadruped, Fast-LIO2 localization and mapping, Hobot-Stereo depth perception, PCT-Planner global planning, EGO-Planner local motion generation, and OpenClaw-orchestrated Qwen3-VL inspection analysis. Experiments in a residential property environment achieved 100\% waypoint reachability, outdoor localization error below 10~cm, local obstacle-response latency below 200~ms, representative hazard-detection rates of 85--95\%, and 99\% success in alarm delivery and structured-report generation. These results validate the deployed navigation and inspection closed loop, while HROS provides an extensible software foundation for memory-augmented, voice-aware, and continuously improvable embodied inspection agents.

Unknown authors · 0 citations
Conference Jul 2026

A Taxonomy of Human-Robot Teamwork Requirements

Autonomous systems are increasingly deployed in safety- and mission-critical domains where humans and robots must operate as a team to complete complex tasks. Existing requirements for Human-Robot teamwork remain fragmented across disparate sources, with no unified framework that addresses complexities of collaborative Human-Robot tasks. We address this gap by presenting a taxonomy of Human-Robot Teamwork (HRT) requirements derived from analysis of (academic and industrial) literature, standards and regulatory guidance. We extracted a construction corpus of 361 requirements from 14 cross-domain sources. Through iterative classification and refinement, we develop a two-level hierarchical taxonomy comprising 6 high-level categories and 21 low-level subcategories that distinguish information provision, relational control, decision support, safety mechanisms, performance monitoring, and foundational system capabilities. We validate the taxonomy through expert evaluation with 5 domain specialists and a utility demonstration on an independently assembled corpus of 448 requirements drawn from 19 sources spanning six HRT domains.

Anastasia Mavridou, Hazel M. Taylor, S. Lozito et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.