We present DiVA, a deeply interactive digital life simulator pioneering a new paradigm for long-term, open-ended interactive experiences within digital character worlds. DiVA's architecture pairs a Multimodal Large Language Model (MLLM) as a router with a meticulously designed stacked video pipeline for seamless, multi...
Cheng Chen, Hao Ouyang, Qiuyu Wang et al.· 0 citations
AnyWorld is proposed, a cross-embodiment world modeling framework that expands a single human interaction into diverse robot-native rollouts without paired human-robot demonstrations and enables independent recomposition of embodiment, viewpoint, and scene factors, allowing a single model to generate many robot-domain...
Cheng Chen, J. Bai, Jiacheng Wei et al.· 3 citations
Image colorization is a fundamental yet challenging task in computer vision, aiming to recover plausible and spatially coherent colors from grayscale images. Recent advancements in diffusion models have enabled significant progress in this field, yet existing methods predominantly rely on multi-step diffusion processes...
Yutong Gao, Ying Zhang, Cong-Yan Lang et al.· IEEE Transactions on Image P...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.