Skip to content

HEROIC: Heterogeneous Evidential Reasoning for Open-Vocabulary Identification and Cross-Robot Collaboration

Sep 2026 · 0 citations · 23 references
Computer Science

TL;DR

HEROIC is a decentralized heterogeneous multi-agent open-vocabulary search coordination framework that requires agents to communicate in natural language only and reaches the target 84% of the time across all 6 scenes, compares to 35-54% for vision-language frontier baselines, frontier-based search, lawnmower, and random-walk running the same perception.

Abstract

Multi-agent heterogeneous air-ground robot teams are attractive for open world search, with applications for reconnaissance, urban search and rescue missions (USAR), disaster response and recovery, and hazardous environments. These two platforms have different failure modes: aerial robots cover ground quickly but cannot resolve small or occluded targets from altitude, while ground robots can identify objects-of-interest, such as people or hazardous objects, at close range but cover less area. Existing language-tasked teams either have roles fixed prior, or have a language model assign them from hand-written capability tags, so the team is unable to know when within a mission an asset is no longer useful. We present HEROIC, a decentralized heterogeneous multi-agent open-vocabulary search coordination framework that requires agents to communicate in natural language only. HEROIC's initial agent role assignment is derived from sensor properties and a scale law to determine whether targets can be detected with a high confidence. From the mission's natural language prompt alone, this law assigns aerial flight altitudes and sweep spacing. When this calculated height falls below the altitude for safe flight, aerial agents re-task themselves from searcher to aerial triage, escort, and route guide for ground agents. Both robots maintain an evidential belief over the search area (bearing rays for positive evidence, a log-odds posterior for negative evidence) and gate any arrival on close-range verification. In full-stack experiments, HEROIC reaches the target 84% of the time across all 6 scenes, compares to 35-54% for vision-language frontier baselines, frontier-based search, lawnmower, and random-walk running the same perception, all while being 2-4x sooner to arrive at the target.

View source

Similar papers

Preprint Sep 2026

AgenticSwarm: Semantic Perception and Adaptive Task Allocation for Heterogeneous Multi-UAV Missions

Multi UAV missions in complex environments require the system to understand both the surrounding scene and the intent of a human operator while maintaining feasible task allocation as mission conditions change. This paper presents AgenticSwarm, an agentic framework for semantic perception and adaptive task allocation i...

Muhammad Ahsan Mustafa, Yasheerah Yaqoot, Faryal Batool et al. · 0 citations
Preprint Sep 2026

Structured World-State Reasoning for Agentic Robotic Search

WorLDS: World-state Observation and Reasoning for Language-guided Discovery and Search, a framework that grounds reasoning in a persistent graph initialized from geospatial priors and updated by perception, is presented.

Finley R. Holt, Luis A. Pabon, J. Alora et al. · 0 citations
#artificial intelligence Open access Aug 2026

Safe Multi-Robot Coordination via VLM–LLM Reasoning and Reachability Analysis

This study presents a centralized safety aware M2M framework for cooperative goal-directed navigation in a heterogeneous mobile robot system composed of a vision-capable robot and a cameraless robotic vehicle that can approve safe motion, trigger conservative replanning or holding behavior, and preserve a strict separa...

Mohamed Dwedar, Ahmad Hafez, Alexander Jesser et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Towards Embodied Air-Ground Cooperative Object Search: Benchmark, Dataset and Agentic Method

Air-Ground Object Search (AGOS) in urban environments is a challenging embodied task, which requires an Unmanned Aerial Vehicle (UAV) and an Unmanned Ground Vehicle (UGV) to jointly search for and verify a specified target vehicle from multi-view visual references. To study this underexplored problem, we introduce AGOS...

Bokyung Yu, Zi-Mo Chen, Jun-Reng Rao et al. · 1 citation · ⚡1
#reinforcement learning Review Open access Sep 2026

From Programmed Execution to Autonomous Reasoning: The LLM-Driven Paradigm Shift in Air-Ground Collaborative Systems

This review critically examines how large language models (LLMs) and multimodal large language models (MLLMs) are reshaping the intelligence paradigm of air-ground heterogeneous collaborative systems between 2023 and 2026 and argues that the core of this transformation is a change in how AGCS understand tasks, coordina...

Zhao-Hui Wang, Yi-Ming Nie, Yan-Xu Hong et al. · 0 citations
Preprint Sep 2026

CoRelNav: Collaborative Relational Navigation for Multi-Robot Spatially Constrained Semantic Navigation

CoRelNav is proposed, whose core is coupling task-conditioned multi-robot exploration with candidate-driven collaborative verification, which reduces redundant search and enables relation hypotheses to be resolved from distributed partial evidence that independent exploration or isolated-view verification can leave amb...

Jin-Yu He, Zi-Hao Mao, Hao-Nan Jin et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.