Skip to content

RoboTalk: Learning Multi-Robot Communication and Coordination from Multimodal Demonstrations

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

RoboTalk is introduced, a synthetic data-generation pipeline and dataset of 7,950 multimodal trajectories spanning 53 mobile-manipulation kitchen tasks for training small VLMs to communicate and coordinate and fine-tuning open-source models on this dataset can reach 77% success on novel held-out tasks.

Abstract

Multi-robot collaboration could enable more efficient and scalable solutions to complex robotic tasks, but collaboration under partial observability remains challenging. Natural-language communication offers a promising approach to coordinating robots under partial observability. However, in decentralized manipulation, jointly learning explicit inter-robot communication and skill-level action selection from multimodal demonstrations remains underexplored for small vision-language models (VLMs) intended for on-device deployment. To address this gap, we introduce RoboTalk, a synthetic data-generation pipeline and dataset of 7,950 multimodal trajectories spanning 53 mobile-manipulation kitchen tasks for training small VLMs to communicate and coordinate. The dataset includes a leader-follower planning protocol, tool calls (perception, manipulation, navigation, and communication), rationale traces, and diversified natural-language communication. Fine-tuning open-source models on our dataset can reach 77% success on novel held-out tasks, a significant improvement over the untuned open source models, which had a success rate of around ~2%.

View source

Similar papers

Preprint Sep 2026

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

MotorMind is introduced, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates, and shows that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchro...

Bing-Xuan Li, Si-Qi Song, Yi-Zhuo Wu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors whil...

Han-Chu Zhou, De-Chen Gao, Hang Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning

We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (S...

Rui-Xiao Xu, W. Kenny, Zhi-Qian Liu et al. · 0 citations
Feb 2025

RobotMover: Learning to Move Large Objects From Human Demonstrations

This work presents RobotMover, a complete learning-based system for large-object manipulation that leverages human–object interaction demonstrations to train robot control policies and achieves strong performance in terms of capability, robustness, and controllability, outperforming both learned and teleoperation basel...

Tian-Yu Li, Joanne Truong, Tsung-Yen Yang et al. · 2 citations
#artificial intelligence Preprint Sep 2026

AeroManip-VLA: Scalable Vision-Language-Action Learning for Aerial Manipulation with RL-Generated Demonstrations

Aerial manipulators extend robotic manipulation into 3D workspaces that are difficult for ground-based robots to access, creating new opportunities for general-purpose manipulation. However, extending Vision-Language-Action (VLA) models to aerial robots introduces distinct challenges due to the tight coupling between m...

Rui Huang, Yan-Lin Mu, Li-Dong Li et al. · 0 citations
Preprint Aug 2026

DREAM: Deployment-Time Demonstration Generation via Real-to-Sim for Scalable Policy Adaptation

DREAM is presented, a framework that generates fine-tuning data for a pretrained VLA from a captured workspace and a language instruction, without requiring a task-specific human demonstration, and whether it can serve as a scalable data-collection system for the deployment workspace.

Makoto Sato, T. Matsushima, Yutaka Matsuo et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Jul 9, 2026

Tiny robot boats build floating structures

MIT researchers developed FloatForm, a swarm of small aquatic robots that snap together like ants forming a raft, assembling into reconfigurable structures on the water.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.