Skip to content
Preprint

EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation

Jul 2026 · 0 citations · 45 references
Computer Science

TL;DR

Experiments on simulated and real-robot manipulation benchmarks demonstrate that EDAR improves downstream policy learning, especially in long-horizon manipulation, highlighting the importance of grounding action representations in executable control structure and environment-conditioned visual change.

Abstract

Learning effective action representations is critical for robotic manipulation, where raw control trajectories are often noisy, redundant, and difficult to model directly. Existing methods mainly encode the structure of the action stream itself, treating the role of actions in the environment as implicit. Yet manipulation is about changing the world: the same action segment can induce different outcomes under different scene contexts, making action semantics inherently environment-dependent. We propose EDAR, an Environment-Dependent Action Representation that grounds action tokens in both executable control structure and expected visual consequences. By coupling motor commands with their environment-conditioned effects, EDAR encourages the learned action space to capture interaction semantics rather than merely command-level patterns. Experiments on simulated and real-robot manipulation benchmarks demonstrate that EDAR improves downstream policy learning, especially in long-horizon manipulation. These results highlight the importance of grounding action representations in executable control structure and environment-conditioned visual change.

View source

Similar papers

Preprint Aug 2026

GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions

GeniWorld is presented, an interactive world model for robots that generalizes robustly across unseen scenarios by explicitly decoupling embodiment kinematics from environmental dynamics, and generates diverse manipulation trajectories within the world model, improving downstream policy performance and robustness in complex environments.

Chenghao Gu, Hanyang Yu, Jingbo Zhang et al. · 1 citation · ⚡1
Preprint Aug 2026

Hydra-0: Action Flow for Generalist World Modeling and Control

We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.

Hongyu Li, Bowen Wen, Xinghao Zhu et al. · 0 citations
Preprint Aug 2026

SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

SLIM (Self-supervised Latent Interaction Model), a compact 0.5B-parameter latent interaction policy, which matches or exceeds representative large-scale VLA and world-action-model baselines with fewer parameters, no additional embodied pretraining, lower inference latency, and substantially lower GPU memory usage.

Jingkai Wang, Zihan Tang, Gu Zhang et al. · 0 citations
Jul 2026

Masked Visual Actions for Unified World Modeling

Masked Visual Actions is introduced, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video that supports inverse modeling by synthesizing robot motion from desired object motion.

Hadi Alzayer, Wenlong Huang, Haonan Chen et al. · 3 citations · ⚡1
Preprint Aug 2026

JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment

JoyAI-RA 0.5 is proposed, a generalist Vision-Language-World-Action framework that couples physical world-dynamics priors with visual semantics and scales manipulation learning across such data via dual action alignment, suggesting that abundant but weakly labeled human experience can be converted into a transferable training signal.

JoyAI-RA Team · 0 citations
Jul 2026

Native Video-Action Pretraining for Generalizable Robot Control

LingBot-VA 2.0 is presented, a video-action foundation model built from the ground up for embodiment, which introduces a semantic visual-action tokenizer, which aligns visual representations with both semantics and actions, improving instruction following and action precision in subsequent policy learning.

Qihang Zhang, Lin Li, Luyao Zhang et al. · 11 citations · ⚡3

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.