Reinforcement Fine-Tuning~(RFT) has emerged as a promising paradigm for improving Vision-Language-Action~(VLA) policies, yet sparse task-level outcomes provide limited credit for intermediate transitions, especially in long-horizon manipulation. A natural approach is to model intermediate task progress and use it as de...
Yun-Peng Qing, Yi-Lun Kong, Si-Xu Lin et al.· 0 citations
Causal Imprint is introduced, which learns future-relevant scene changes from training-only future supervision and provides predictive representations directly to the action expert without future-video rollout at inference.
Xing-Yu Miao, Zi-Zun Li, Bao-Le Fang et al.· 0 citations
Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal interfaces, asynchronous multi-...
Ji-Song Cai, Yao Mu, Gan-Lin Yang et al.· 0 citations
By constraining strategic exploration through a pretrained motion decoder, RoboStriker substantially reduces the catastrophic balance failures observed in raw action-space methods and achieves superior tactical performance in both competitive win rates and striking efficiency.
Kang-Ning Yin, Kaige Liu, Zhe Cao et al.· 0 citations
This work revisits the scaling recipe for BFMs and demonstrates that substantial performance gains can be achieved through the coordination of three core components: the learning paradigm of motion tracking that reformulates diverse humanoid control problems as the reproduction of integrated whole-body behaviors in the...