World-action models (WAMs) jointly generate future world states and actions through iterative denoising, using shared weights to process heterogeneous semantic streams of video, proprioceptive, and action tokens. Quantization reduces inference cost, but comparable numerical errors in different streams can have markedly...
Yun-Han Wang, Hao-Dong Wang, Zhi-Ming Liu et al.· 0 citations
World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and actions. This capability comes at a substantial inference cost, as dense future-frame tokens are repeatedly processed during denoising. Prior methods address this by toke...
Xin-Ling Xie, Hao-Dong Wang, Jia-Zhi Mi et al.· 0 citations
HATCH (Hint-Annealed Self-Teaching), an online single-policy framework that learns from both generating and using its own hints to improve reasoning without assistance, is proposed and gradient projection is used to remove the opposing component of hint-generation updates.
Zi-Le Wang, Zi-Jian Li, Hao-Dong Wang et al.· 0 citations
Generative world models aim to predict future states conditioned on actions, where action controllability is fundamental for reliable dynamics modeling. While recent efforts leverage simulator-generated data to enhance this capability, existing training pipelines face two fundamental limitations. First, static offline...
Yi-Kun Miao, Fang-Qi Zhu, Quan-Xin Shou et al.· 2 citations
Large language models (LLMs) are increasingly used for everyday assistance, yet existing benchmarks only partially reflect the requests users naturally make in practice. Real-world requests are often open-ended, casually specified, and context-dependent, requiring models not only to follow explicit instructions but als...
Yong Peng, Qing-Shui Gu, Li-Ya Zhu et al.· 0 citations
This work introduces Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads.
Zining Huang, Haoran Que, Hongxia Zeng et al.· 2 citations
This work systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains and establishes StartupBench as an empirical measure of progress toward E2E completions o...
Li-Ya Zhu, Xin Ma, Tao Liu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.