This work presents Enfold, which transfers this computation that constructs a future into a representation predicted from the current visual context and language instruction, and recast a world generator as a source of predictive control representations if its internal structure can be enfolded into the present.
Wei-Li Zeng, Yi-Tong Xing, Fu-Long Liu et al.· 1 citation
A unified spatiotemporally decoupled framework named DeMoDiff is proposed, which jointly redesigns representation and architecture and incorporates spatial-temporal masking and attention mechanisms into an autoregressive diffusion generator, achieving both generative capability and controllable editability.
Chengqun Yang, Liang Xu, Yanping Li et al.· 0 citations
This work proposes a lightweight granularity-aware model anchored at a frozen standard-caption-aligned retrieval model that improves mixed-granularity retrieval without compromising standard-caption performance, and believes that its MRBench provides a comprehensive testbed for advancing motion-language alignment evaluation.
Fulong Liu, Liang Xu, Chengqun Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.