Tool-augmented vision-language models increasingly"think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind tests, gain decompositions, and attention analyses has shown that returned images contribute little, raising the question: if pixels do not car...
Jiahao Shao, Yuanbo Yang, Yi-Yi Liao et al.· 3 citations
This work presents Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance, and proposes an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8....
Jia-Min Zhou, Qihang Zhang, Gangwei Xu et al.· 8 citations
Image2Sim is introduced, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences and serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale.
LingBot-VA 2.0 is presented, a video-action foundation model built from the ground up for embodiment, which introduces a semantic visual-action tokenizer, which aligns visual representations with both semantics and actions, improving instruction following and action precision in subsequent policy learning.
Qihang Zhang, Lin Li, Luyao Zhang et al.· arXiv.org· 11 citations· ⚡3
This work proposes masked boundary modeling, a self-supervised paradigm that dynamically learns sub-pixel boundary representations and subsequently leverages the discovered boundary-bearing tokens as masked targets to facilitate dense visual token learning.
Zelin Fu, Bin Tan, Chang Sun et al.· 3 citations· ⚡2
The integration of an agentic harness within the domain of world modeling is pioneered, wherein a pilot agent is tasked with planning and executing character behaviors, while a director agent is responsible for synthesizing novel environmental elements as the scene progresses.
LingBot-Video is presented, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence, and is contributed as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.
Shuailei Ma, Jiaqi Liao, Xinyang Wang et al.· 7 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.