Skip to content

Author

Yinghao Xu

We have 8 of 75 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Thinking With Tools, Not With Pixels: Tool Calls as Text Scaffolds for Visual Reasoning

Tool-augmented vision-language models increasingly"think with images": they call crop, zoom, or code tools and reason over the returned pixels. However, recent work using blind tests, gain decompositions, and attention analyses has shown that returned images contribute little, raising the question: if pixels do not car...

Jiahao Shao, Yuanbo Yang, Yi-Yi Liao et al. · 3 citations
Preprint Aug 2026

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

This work presents Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance, and proposes an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8....

Jia-Min Zhou, Qihang Zhang, Gangwei Xu et al. · 8 citations
Jul 2026

Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator

Image2Sim is introduced, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences and serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale.

Zi-Han Wang, Seungjun Lee, Yinghao Xu et al. · 0 citations
Preprint Aug 2026

4DAnyone: Create Anyone in 4D from a Casual Monocular Video

4DAnyone outperforms prior methods in both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization.

Yudong Jin, Tao Xie, Qihang Zhang et al. · 0 citations
Jul 2026

Native Video-Action Pretraining for Generalizable Robot Control

LingBot-VA 2.0 is presented, a video-action foundation model built from the ground up for embodiment, which introduces a semantic visual-action tokenizer, which aligns visual representations with both semantics and actions, improving instruction following and action precision in subsequent policy learning.

Qihang Zhang, Lin Li, Luyao Zhang et al. · 11 citations · ⚡3
Preprint Jul 2026

Vision Pretraining for Dense Spatial Perception

This work proposes masked boundary modeling, a self-supervised paradigm that dynamically learns sub-pixel boundary representations and subsequently leverages the discovered boundary-bearing tokens as masked targets to facilitate dense visual token learning.

Zelin Fu, Bin Tan, Chang Sun et al. · 3 citations · ⚡2
Preprint Jul 2026

Infinite Worlds with Versatile Interactions

The integration of an agentic harness within the domain of world modeling is pioneered, wherein a pilot agent is tasked with planning and executing character behaviors, while a director agent is responsible for synthesizing novel environmental elements as the scene progresses.

Zelin Gao, Qiuyu Wang, Jiapeng Zhu et al. · 12 citations · ⚡4
Preprint Jul 2026

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

LingBot-Video is presented, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence, and is contributed as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.

Shuailei Ma, Jiaqi Liao, Xinyang Wang et al. · 7 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.