This work presents Enfold, which transfers this computation that constructs a future into a representation predicted from the current visual context and language instruction, and recast a world generator as a source of predictive control representations if its internal structure can be enfolded into the present.
Wei-Li Zeng, Yi-Tong Xing, Fu-Long Liu et al.· 1 citation
InternVLA-A1.5 is presented, which builds the policy on a native VLM backbone that keeps training on VQA and subtask prediction, and attaches a lightweight unified expert for continuous action generation, and achieves the best overall results on all six simulation benchmarks.
This work presents REAL, an agentic framework for open-world mobile manipulation, which establishes sim-to-real-consistent environment APIs without oracle perception and integrates a simulated user to enable human-in-the-loop interaction.
Boyu Mi, Mengchen Ma, Yifei Yao et al.· arXiv.org· 2 citations
GaussianWAM is proposed, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field and improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation.
Zijian Zhang, Yuqing Jiang, Weitao Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.