Tactile sensing provides essential contact information for robotic manipulation, yet incorporating it into pretrained vision-language-action (VLA) models remains challenging. A common concern is that simply introducing touch during task-specific fine-tuning may fail to bridge the cross-modal gap, yielding limited gains...
Chen Yang, Lin-Zhe Shi, Chang Jie Wu et al.· 0 citations
Results support the effectiveness of caption memory for episodic reasoning over long egocentric video in wearable assistants with bounded frame budgets, growing visual-token costs, and long-context retrieval failures.
Dingli Liang, Yi Xie, Yu-Kai Huang et al.· 0 citations
This work proposes Caption-once, Frames-onDemand (CFD), a budget-aware edge-cloud agentic framework that turns visual access into a first-class, query-conditioned cost, capping per-query frame consumption regardless of video length.
Wei-Tong Cai, Hang Zhang, Yu-Kai Huang et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.