HandWorld: Hand-Centric Unified Video Action Generation
This work introduces HandWorld, a unified generative framework that focuses on hand-object interaction and jointly models ego-centric videos and hand actions and learns shared cross-domain conditions through a dual-branch condition network that integrates information from both video and action domains.