Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?
This work proposes a new learning paradigm that combines hand-object masked training, which enables robust reasoning from partial hand or object observations, and an HOI-dynamics-aware decoder that explicitly learns hand- and object-centric embeddings through auxiliary predictions of their locations and semantics, enhancing sensitivity to both cues.