A unified latent-space framework for image and video diffusion models that achieves the sota performance among various metrics and further improves optimization stability and achieves the highest VBench quality, semantic, and total scores among the evaluated methods.
Rui Li, Yuan-Zhi Liang, Ke-Chun Hao et al.· 0 citations
LLMODE is proposed, a token-efficient framework for irregular spatio-temporal forecasting with a frozen LLM backbone that shows competitive overall performance, with clearer advantages under sparse or dynamically complex irregular sampling.
Di Zhang, Jing-Yang Zhang, Zi-Qian Wang et al.· 0 citations
This work proposes a co-evolution roadmap for physical intelligence centered on theembodied brain, a long-term model target for integrating multimodal context, comparing candidate interventions, and issuing state-transition or capability requests rather than direct actuator commands.
Yuanzhi Liang, Xufeng Zhan, Haibin Huang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.