Tool-augmented vision-language models increasingly"think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind tests, gain decompositions, and attention analyses has shown that returned images contribute little, raising the question: if pixels do not car...
Jiahao Shao, Yuanbo Yang, Yi-Yi Liao et al.· 3 citations
LingBot-VA 2.0 is presented, a video-action foundation model built from the ground up for embodiment, which introduces a semantic visual-action tokenizer, which aligns visual representations with both semantics and actions, improving instruction following and action precision in subsequent policy learning.
Qihang Zhang, Lin Li, Luyao Zhang et al.· arXiv.org· 11 citations· ⚡3
LingBot-Video is presented, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence, and is contributed as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.
Shuailei Ma, Jiaqi Liao, Xinyang Wang et al.· 7 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.