We present DiVA, a deeply interactive digital life simulator pioneering a new paradigm for long-term, open-ended interactive experiences within digital character worlds. DiVA's architecture pairs a Multimodal Large Language Model (MLLM) as a router with a meticulously designed stacked video pipeline for seamless, multi...
Cheng Chen, Hao Ouyang, Qiuyu Wang et al.· 0 citations
CameraAnything is introduced, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic camera parameters, and a scalable synthetic pipeline is developed that constructs diverse dynamic scenes through structured multi-camera recording and generates synchr...
Yixuan Li, Yanhong Zeng, K. Cheng et al.· arXiv.org· 1 citation
LingBot-Video is presented, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence, and is contributed as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.
Shuailei Ma, Jiaqi Liao, Xinyang Wang et al.· 7 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.