Skip to content

TacSushi: Tactile-Grounded World-Action Modeling for Dexterous Sushi Manipulation

Sep 2026 · 0 citations · 34 references
Computer Science

TL;DR

Comparisons support complementary benefits of feature-wise gated tactile fusion and training-only predictive supervision of TacSushi, a tactile-grounded, Cosmos3-based world-action policy that learns from recorded future consequences while acting on current observations.

Abstract

Dexterous food manipulation requires control under deformation, occlusion, and uncertain contact. We present TacSushi, a tactile-grounded, Cosmos3-based world-action policy that learns from recorded future consequences while acting on current observations. The backbone encodes current RGB, language, and hand state, and feature-wise gated fusion incorporates fingertip tactile features into the action representation. During training, a decoder conditioned on demonstrated action chunks predicts logged future visual observations, task progress, relative contact risk, and tactile summaries; this decoder is removed at deployment. Failed trials provide consequence supervision, but their actions are excluded from imitation. We train TacSushi on 340 successful and 50 failed real-robot trials and compare six methods in 600 separate rollouts across three in-distribution tasks and two out-of-distribution ingredient variants. To assess food quality beyond a single geometric threshold, we score terminal outcomes using an anchored visual-quality protocol that equally weights five human ratings and three vision-language-model ratings per rollout. Full TacSushi achieves 68.3% average in-distribution success and 37.5% out-of-distribution success, compared with 36.7%/10.0% without future-consequence supervision and 25.0%/17.5% with direct tactile concatenation in place of gated fusion. These comparisons support complementary benefits of feature-wise gated tactile fusion and training-only predictive supervision.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation

Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a...

Hao-Ran Yuan, Ze-Kai Wang, Boning Shao et al. · 0 citations
Preprint Sep 2026

UniDex-ViTac: Learning Unified Visuo-Tactile Dexterous Manipulation Policy from Human Video Data

Human videos provide demonstrations of dexterous manipulation but lack robot-executable actions and tactile measurements. We present UniDex-ViTac, a framework that uses human-video-guided simulation to generate robot demonstrations paired with fingertip contact observations for training a deployable visuo-tactile polic...

Hyesung Lee, Si-Hwan Heo, Sungwook Yang · 0 citations
Preprint Sep 2026

TacBPM: A Tactile-conditioned Behavior Prior Model for Dexterous Reorientation

Dexterous in-hand manipulation requires policies that coordinate high-DoF hand joints through intermittent, contact-rich interaction. Beyond target-orientation tracking, such policies must discover finger gaits that preserve object stability while adapting to geometry, anisotropy, pose, contact, and sensing changes. We...

Jie Yin, Wan-Li Xing, Ze-Yuan Zhao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation

Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated b...

Wen-Qiao Li, Qian-You Zhao, Jia-Wen Hao et al. · 0 citations
Preprint Sep 2026

DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation

Learning predictive models of contact-rich dexterous manipulation requires dense tactile interaction, but such data are costly to scale on real robots and remain tied to embodiment-specific sensors. We introduce DexTouch-WM, an action-conditioned world model that learns from scalable human touch to jointly predict futu...

Yan-Jiao Qin, Yue Chen, Wen-Wei Lin et al. · 0 citations
Preprint Sep 2026

TACIT: Tactile Contact Supervision for Spatial Attention in Dexterous Manipulation

Visuomotor policies trained from a few demonstrations may reproduce demonstrated trajectories without reliably following changes in object position. Existing approaches with explicit attention typically obtain spatial priors from human annotation or visual models. We introduce TACIT (tactile contact informs attention),...

Yan-Hou Lai, Fu-Cai Zhu, Rui-Qiang Wang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.