V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
P predictive visual latents are established as a foundation for effective WAM learning from task-specific demonstrations and for transferring future-modeling knowledge acquired from broader in-the-wild videos.