World-action models (WAMs) jointly generate future world states and actions through iterative denoising, using shared weights to process heterogeneous semantic streams of video, proprioceptive, and action tokens. Quantization reduces inference cost, but comparable numerical errors in different streams can have markedly...
Yun-Han Wang, Hao-Dong Wang, Zhi-Ming Liu et al.· 0 citations
World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and actions. This capability comes at a substantial inference cost, as dense future-frame tokens are repeatedly processed during denoising. Prior methods address this by toke...
Xin-Ling Xie, Hao-Dong Wang, Jia-Zhi Mi et al.· 0 citations
HATCH (Hint-Annealed Self-Teaching), an online single-policy framework that learns from both generating and using its own hints to improve reasoning without assistance, is proposed and gradient projection is used to remove the opposing component of hint-generation updates.
Zi-Le Wang, Zi-Jian Li, Hao-Dong Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.