Skip to content
Preprint

WM-Craftnet: World Synesthesia Model for Generalizable and Robust Dexterous In-Hand Manipulation

Sep 2026 · 2 citations · 24 references
Computer Science

TL;DR

WM-Craftnet is proposed, a world-model-conditioned framework that learns compact action-conditioned latent dynamics from proprioception, depth, tactile sensing, and actions, supervised by multimodal reconstruction and reward prediction.

Abstract

Generalizable and robust dexterous in-hand manipulation requires a policy to infer object pose, geometry, contact, and potential slip from partial and noisy observations. Although recent tactile and visuotactile RL methods achieve strong in-hand rotation in controlled settings, their robustness often degrades under pose shifts, force disturbances, and object variation. We propose WM-Craftnet, a world-model-conditioned framework that learns compact action-conditioned latent dynamics from proprioception, depth, tactile sensing, and actions, supervised by multimodal reconstruction and reward prediction. Rather than using the world model for latent imagination or policy optimization, WM-Craftnet uses the learned World Synesthesia Model (WSM) as recurrent task context for an asymmetric actor--critic policy. Importantly, WSM is trained to reconstruct clean depth targets from noisy depth inputs, providing a denoised geometric state for real-robot deployment. Ablations over recurrent baselines, auxiliary heads, tactile masking, and WSM modality heads show that predictive world modeling, clean-depth supervision, and tactile contact cues all shape the learned state. A WSM pretrained on nine \(z\)-axis objects serves as a reusable prior for \(49\)-object downstream policy learning. This context improves multi-object rotation, with quantitative and qualitative evidence for unseen-object, perturbation-recovery, and sim-to-real transfer.

View source

Similar papers

Preprint Aug 2026

DreamMimic: Learning Visuomotor Whole-Body Loco-Manipulation via World Model

A framework that distills privileged teacher policies into vision-based humanoid controllers via world-model-assisted distillation, and introduces Performance-Conditioned Guidance (PCG), a reward-driven adaptive distillation schedule that computes performance scores for both teacher and student to dynamically balance g...

Jie Yin, Xing-Yu Lai · 1 citation
#artificial intelligence Preprint Sep 2026

TacSushi: Tactile-Grounded World-Action Modeling for Dexterous Sushi Manipulation

Comparisons support complementary benefits of feature-wise gated tactile fusion and training-only predictive supervision of TacSushi, a tactile-grounded, Cosmos3-based world-action policy that learns from recorded future consequences while acting on current observations.

Hao-Di Hu, Kaen Kogashi, T. Koike-Akino · 0 citations
Preprint Sep 2026

UniDex-ViTac: Learning Unified Visuo-Tactile Dexterous Manipulation Policy from Human Video Data

Human videos provide demonstrations of dexterous manipulation but lack robot-executable actions and tactile measurements. We present UniDex-ViTac, a framework that uses human-video-guided simulation to generate robot demonstrations paired with fingertip contact observations for training a deployable visuo-tactile polic...

Hyesung Lee, Si-Hwan Heo, Sungwook Yang · 0 citations
#artificial intelligence Preprint Sep 2026

DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation

Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a...

Hao-Ran Yuan, Ze-Kai Wang, Boning Shao et al. · 0 citations
#machine learning Preprint Sep 2026

Dexterous Tactile World Model

World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future...

Zi-Yao Zeng, Xiatao Sun, Hao Wang et al. · 0 citations
Preprint Aug 2026

Robust Bimanual Vision-Language-Action Models via Embarrassingly Simple Modality Masking

The Modality Masking Mechanism (M3), an embarrassingly simple, training-only strategy that requires no architectural changes or large-scale robot pretraining, is introduced to improve robustness of query-based VLA policies for bimanual manipulation.

Dong-Zhou Cheng, Ziang Li, Yixiao Zhou et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.