Towards Predictive, Aligned, and Scalable Robot Learning
Lumo-2 is introduced, a latent world-action model that generates actions by reasoning over world dynamics in latent space that consistently outperforms strong vision-language-action and world-action model baselines, with gains on challenging real-world tasks requiring temporal reasoning, physical understanding, or high control complexity.