Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfol...
Yu-Bo Zhu, Ya-Wen Shao, Zi-Yun Dai et al.· 0 citations
VIVAS is proposed, a framework built upon the unified token space paradigm, which introduces a dense-structural-semantic vision tokenizer, which expands the textual vocabulary into a unified vision-language vocabulary by incorporating a visual vocabulary.
Zhe-Han Kan, Yu-Bo Zhu, Xing-Hua Jiang et al.· 0 citations
UVU effectively synergizes pixel-level visual perception with semantic-level visual understanding, internalizing visual reconstruction capabilities and unlocking the facilitative role of visual supervision in enhancing understanding in the pre-training stage.
Zhe-Han Kan, Xing-Hua Jiang, Yu-Bo Zhu et al.· 0 citations
VGAU-Diag is introduced, a fine-grained evaluation framework for vision generation-assisted understanding that stratifies samples by difficulty, enables unified evaluation of multiple reasoning paradigms, and uses Oracle-Ass Reference Protocols.
Yu-Bo Zhu, Zhe-Han Kan, Jing-Yi Yang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.