Skip to content
Preprint

MoPA: Coordinated Mobile Manipulation via Subsystem-Specific Perception Alignment

Sep 2026 · 2 citations · 42 references
Computer Science

TL;DR

MoPA is presented, a framework that aligns perceptual conditioning with mobility and manipulation while preserving coordination at the action level, and achieves state-of-the-art performance across all three task suites.

Abstract

Mobile manipulation requires perceptual evidence at different spatial scales for base motion and arm control, while the two action modalities remain kinematically coupled. Existing policies often employ specialized action generation for different subsystems but condition heterogeneous action branches on a shared perceptual representation, leaving subsystem-specific perception-action correspondence implicit. We present MoPA, a framework that aligns perceptual conditioning with mobility and manipulation while preserving coordination at the action level. Dual Perceptual Streams employ two mutually masked query banks to extract separate perceptual representations from a shared vision-language context. Perception2Action Adaptation jointly updates each query bank and its corresponding action stream at every layer of a structured Mixture-of-Transformers decoder, while enabling information exchange between the two action streams. Coupled conditional flow matching learns a joint vector field for coordinated generation of both action chunks. On the ManiSkill-HAB benchmark, MoPA achieves state-of-the-art performance across all three task suites. Across four real-world tasks, MoPA achieves a mean full-task success rate of 76.3%, outperforming the best baseline by 12.5 percentage points. Ablation studies and further analyses validate the effectiveness of the proposed design. Website is available at: https://mopa-policy.github.io/.

View source

Similar papers

Preprint Sep 2026

MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining

Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing appro...

Qi-Wei Liang, Guang-Yu Chen, Shao-Long Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MAVP: Map-Aware Visuomotor Policies for Mobile Manipulation

Successful mobile manipulation requires coordinated base and arm motion while maintaining accurate spatial positioning. However, demonstration-trained policies can struggle to realise the intended base motion reliably, leading to spatial misalignment and subsequent manipulation failures. We present MAVP (Map-Aware Visu...

Jin-He Tang, Rui Dai, Wei-Ming Zhi · 0 citations
Preprint Aug 2026

TONAV: Task-Oriented Navigation and Action-Velocity Chunk Learning for Articulated Object Quadrupedal Mobile Manipulation

Quadruped mobile manipulation requires two tightly coupled capabilities: reaching manipulation-ready configurations and maintaining stable contact throughout articulated-object interaction. However, existing methods often terminate navigation near the target, leaving a gap between reachability and manipulation readines...

Hao-Ran Lin, Mingyu Yang, Pengfei Qi et al. · 0 citations
Preprint Sep 2026

DualWAM: Dual-System World Action Models for Asynchronous Global Planning and Local Refinement

World Action Models (WAMs) jointly generate robot actions and predict future world states, transferring priors from video pretraining to robot control. However, future visual prediction is computationally expensive, so existing WAMs often rely on long action chunks to amortize inference cost across control steps, at th...

Yi-Xin Zheng, Jiangran Lyu, Yun-Tian Deng et al. · 0 citations
Preprint Aug 2026

StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models

StreamPI is proposed, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters and seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference.

Zhe Liu, Jinghua Hou, Yuxiang Lu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.