Skip to content
Preprint

EgoTac: In-the-wild Tactile Prediction from Egocentric Vision

Aug 2026 · 1 citation · ⚡ 1 influential · 48 references
Computer Science

TL;DR

Overall, EgoTac provides a scalable pathway to extract tactile priors from egocentric human videos, enabling broadly applicable tactile-aware robot learning.

Abstract

Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot learning lacks tactile information. Directly collecting large-scale tactile data is challenging due to sensor limitations, while human video data is abundant, contact-rich, and easily scalable. This motivates a natural question: can tactile signals be inferred purely from vision? To address this, we introduce EgoTac, a generalizable model that predicts rich tactile information directly from egocentric human videos. EgoTac is trained on a unified corpus of over 5.7M image-tactile pairs, covering both continuous force measurements and binary contacts. By learning from this diverse dataset, EgoTac captures nuanced touch dynamics across varied interactions. Experiments demonstrate strong performance: in-domain prediction achieves an average force error below 0.06N. On out-of-domain contact prediction benchmarks, EgoTac consistently outperforms the state-of-the-art contact estimator. It also captures the rise and fall patterns of real tactile data and enables zero-shot predictions on unconstrained real-world videos. Scaling analyses further reveal that both data diversity and volume improve performance steadily. Overall, EgoTac provides a scalable pathway to extract tactile priors from egocentric human videos, enabling broadly applicable tactile-aware robot learning.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

TouchSight: Bare-Handed Tactile Prediction from Egocentric Video via Generative Visual Augmentation

Tactile signals provide direct contact and force measurements that are essential for understanding physical interactions and enabling dexterous robotic manipulation. However, tactile sensing requires direct measurement at contact interfaces, making large-scale data collection reliant on intrusive, costly, and restricti...

Dan-Yan Zhou, Jin-Xuan Lu, Jia-Wei Lin et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DexTacWAM: A Visuo-Tactile World-Action Model for Dexterous Manipulation

Dexterous manipulation depends on contact dynamics that are often only partially observable from vision. Recent World-Action Models (WAMs) couple predictive video world modeling with action generation, but remain largely vision-centric and therefore cannot directly model these contact dynamics. We present DexTacWAM, a...

Hao-Ran Yuan, Ze-Kai Wang, Boning Shao et al. · 0 citations
Preprint Sep 2026

TACTIC: Understanding Tactile Encoders and Conditioning for Contact-rich Robot Manipulation Policies

Tactile information is essential for contact-rich manipulation tasks in robotics. Vision-based tactile sensors make it particularly easy to design end-to-end manipulation policies with tactile sensing, as they enable the use of existing encoders from computer vision. However, this has led to a huge variety of architect...

Seongjin Bien, Débora Oliveira Makowski, Carlo Kneissl et al. · 0 citations
Preprint Sep 2026

STAR: Sparse Tactile Representation Learning in Vision-Tactile-Language-Action Models for Dexterous Manipulation

Dexterous manipulation requires coordinated multi-finger control and effective tactile feedback, yet learning these capabilities remains challenging due to the lack of large-scale real-world data and the difficulty of extracting effective representations from sparse tactile signals. We build a robot platform and teleop...

Xiang-Cheng Liu, Tian-Hao Wu, Le Zheng et al. · 1 citation
#machine learning Preprint Sep 2026

Dexterous Tactile World Model

World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future...

Zi-Yao Zeng, Xiatao Sun, Hao Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation

Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated b...

Wen-Qiao Li, Qian-You Zhao, Jia-Wen Hao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.