We present DiVA, a deeply interactive digital life simulator pioneering a new paradigm for long-term, open-ended interactive experiences within digital character worlds. DiVA's architecture pairs a Multimodal Large Language Model (MLLM) as a router with a meticulously designed stacked video pipeline for seamless, multi...
Cheng Chen, Hao Ouyang, Qiuyu Wang et al.· 0 citations
CameraAnything is introduced, the first unified framework for camera controlled video editing that enables joint control of both intrinsic and extrinsic camera parameters, and a scalable synthetic pipeline is developed that constructs diverse dynamic scenes through structured multi-camera recording and generates synchr...
Yixuan Li, Yanhong Zeng, K. Cheng et al.· arXiv.org· 1 citation
LingBot-VA 2.0 is presented, a video-action foundation model built from the ground up for embodiment, which introduces a semantic visual-action tokenizer, which aligns visual representations with both semantics and actions, improving instruction following and action precision in subsequent policy learning.
Qihang Zhang, Lin Li, Luyao Zhang et al.· arXiv.org· 11 citations· ⚡3
The integration of an agentic harness within the domain of world modeling is pioneered, wherein a pilot agent is tasked with planning and executing character behaviors, while a director agent is responsible for synthesizing novel environmental elements as the scene progresses.
LingBot-Video is presented, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence, and is contributed as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.
Shuailei Ma, Jiaqi Liao, Xinyang Wang et al.· 7 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.