WorldExam is introduced, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity, which supports unified evaluation of camera-, action-, and language-driven model paradigms.
Yuxue Yang, Shu-Yao Shang, Jiahe Wang et al.· 2 citations
Recent advances in humanoid robotics have produced diverse skills through reinforcement learning, motion imitation, and generative modeling. Yet these capabilities remain siloed because they are built around incompatible representations, interfaces, and controllers. We present CHOREO, a framework for training-free comp...
Zi-Yi Sun, Jing-Wen Chen, Yu-Xin Wang et al.· 0 citations
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors. Mot...