Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI
This survey presents a critical review of VLMs for egocentric video understanding, tracing the progression from conventional recognition architectures to multimodal foundation models and embodied systems, and examines how first-person perception and multimodal foundation models support wearable assistance, robot skill learning, human-to-robot transfer, and embodied decision making.
M. Zamani, Fatemeh Ziaeetabar
· 0 citations