RoboTalk is introduced, a synthetic data-generation pipeline and dataset of 7,950 multimodal trajectories spanning 53 mobile-manipulation kitchen tasks for training small VLMs to communicate and coordinate and fine-tuning open-source models on this dataset can reach 77% success on novel held-out tasks.
Abstract
Multi-robot collaboration could enable more efficient and scalable solutions to complex robotic tasks, but collaboration under partial observability remains challenging. Natural-language communication offers a promising approach to coordinating robots under partial observability. However, in decentralized manipulation, jointly learning explicit inter-robot communication and skill-level action selection from multimodal demonstrations remains underexplored for small vision-language models (VLMs) intended for on-device deployment. To address this gap, we introduce RoboTalk, a synthetic data-generation pipeline and dataset of 7,950 multimodal trajectories spanning 53 mobile-manipulation kitchen tasks for training small VLMs to communicate and coordinate. The dataset includes a leader-follower planning protocol, tool calls (perception, manipulation, navigation, and communication), rationale traces, and diversified natural-language communication. Fine-tuning open-source models on our dataset can reach 77% success on novel held-out tasks, a significant improvement over the untuned open source models, which had a success rate of around ~2%.
MotorMind is introduced, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates, and shows that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchro...
Bing-Xuan Li, Si-Qi Song, Yi-Zhuo Wu et al.· 0 citations
Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progress has focused on single-robot settings. Extending these capabilities to multi-robot systems remains challenging because robots must coordinate long-horizon behaviors whil...
Han-Chu Zhou, De-Chen Gao, Hang Wang et al.· 0 citations
We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (S...
Rui-Xiao Xu, W. Kenny, Zhi-Qian Liu et al.· 0 citations
This work presents RobotMover, a complete learning-based system for large-object manipulation that leverages human–object interaction demonstrations to train robot control policies and achieves strong performance in terms of capability, robustness, and controllability, outperforming both learned and teleoperation basel...
Tian-Yu Li, Joanne Truong, Tsung-Yen Yang et al.· IEEE Transactions on robotic...· 2 citations
Aerial manipulators extend robotic manipulation into 3D workspaces that are difficult for ground-based robots to access, creating new opportunities for general-purpose manipulation. However, extending Vision-Language-Action (VLA) models to aerial robots introduces distinct challenges due to the tight coupling between m...
Rui Huang, Yan-Lin Mu, Li-Dong Li et al.· 0 citations
DREAM is presented, a framework that generates fine-tuning data for a pretrained VLA from a captured workspace and a language instruction, without requiring a task-specific human demonstration, and whether it can serve as a scalable data-collection system for the deployment workspace.
Makoto Sato, T. Matsushima, Yutaka Matsuo et al.· 1 citation
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduJul 9, 2026
MIT researchers developed FloatForm, a swarm of small aquatic robots that snap together like ants forming a raft, assembling into reconfigurable structures on the water.
MIT News · Artificial Intelligence· news.mit.eduJun 26, 2026
To help robots do chores in places like homes and factories, a new approach from MIT uses one language model to clarify users’ instructions, then another to ignore irrelevant info.