Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
This work presents Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance, and proposes an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8....