Towards Generalizable Visuotactile Policies for Contact-Rich Robotic Manipulation
Contact-rich robotic manipulation requires a robot to reason not only about the geometry and semantics of its environment, but also about physical interactions that become observable only through contact. While vision provides rich global information, it can be unreliable during close interaction due to occlusion, viewpoint changes, and limited access to contact dynamics. Tactile sensing provides complementary local information about physical interaction, yet existing visuotactile policies often require large amounts of task-specific data and may generalize poorly across spatial configurations, objects, tasks, and sensing conditions. This thesis investigates generalizable visuotactile policies for data-efficient contact-rich robotic manipulation. The central goal is to develop learning methods that effectively integrate visual and tactile information while preserving task-relevant structure and enabling robust transfer beyond the demonstrated training conditions. The research explores several complementary directions, including structured and equivariant multimodal representations, visuotactile fusion, generative action policies, and efficient adaptation of pretrained visuomotor or vision-language-action models using tactile feedback. Building upon our work on equivariant visuotactile diffusion policies, the thesis will study how tactile information can improve spatial robustness, contact-aware action generation, and task-level generalization under limited demonstrations. The proposed methods will be evaluated on real-world robotic manipulation tasks involving challenging contact, occlusion, and variations in object pose, task configuration, and interaction conditions. Ultimately, this research aims to establish more general and reusable approaches for incorporating tactile feedback into learning-based robotic manipulation.