Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use h...
Yi-Ze Liu, Huang Huang, Yi-Ning Hong et al.· 0 citations
A robot may lose sight of an object it must later retrieve, need to recall what a person demonstrated earlier, or track which steps of a task it has already completed. Current vision-language-action (VLA) policies often fail once the information needed for action disappears from the current observation, making memory c...
Yi-Ze Liu, Ke Wang, Mac Schwager et al.· 0 citations
Reinforcement learning for multi-turn search reasoning typically relies on terminal outcome rewards, which cannot distinguish useful, redundant, and harmful intermediate interactions. We propose LOTAPO , a self-generated process-supervision method based on backward leave-one-turn attribution. For each search turn, LOTA...
Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Cognitive science identifies these relations as foundational to spatial understanding, yet foundation-model evaluations largely focus on metri...
Yun-Fei Ge, An-Bang Liu, Qineng Wang et al.· 0 citations
DemoMimic (Dexterous Motion Mimic), a policy that manipulates objects by focusing on their geometry local to the contact points to improve sim-to-real consistency, yielding a single real-world policy that transfers across objects of varying shape, scale, mass, and friction wherever local contact structure is preserved.
TrAct is proposed, a world-model-based robot decision-making framework that uses visual tracks as an intermediate interface between control and prediction, enabling more accurate world modeling and stronger robot generalization.
Zhi-Hang Cao, Howard Ji, Kevin Zhang et al.· 3 citations
FL-MAESTRO is proposed, a multi-agent orchestrator that makes the joint runtime FL decision directly through three specialist LLM agents, one per decision dimension, and matches the accuracy of the strongest energy-aware baseline while cutting wasted round energy from over a third to near zero.
Jiajun Wu, Zirui Wang, Jiayu Zhou et al.· 0 citations
WebWorld is presented, the interface that lets a VLM prior interact with this browser-as-world-model autonomously and decides which interactions become supervision and reaches the level of strong frontier systems such as Kimi-K2.6 and GPT-5.4 on interactive HTML generation.
Jia-Jun Wu, Jian Yang, Ya-Xin Du et al.· 0 citations
Masked Visual Actions is introduced, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video that supports inverse modeling by synthesizing robot motion from desired object motion.
Hadi Alzayer, Wenlong Huang, Haonan Chen et al.· arXiv.org· 3 citations· ⚡1
CyberFactory is introduced, a unified open-source framework that connects data construction, trajectory synthesis, and model training across proof-of-concept (PoC) generation, vulnerability patching, and cybersecurity question answering (CyberQA).
Jian Yang, Haau-Sing Li, Shawn Guo et al.· 0 citations
This work proposes APIVOT, a VLM-based planner that adaptively interleaves language and visual thoughts for long-horizon planning that outperforms general-purpose VLMs and prior planning frameworks, achieving the largest gains in spatially constrained settings.
This work presents LoopCoder pre-trained on 12T+ code and general tokens, along with LoopCoder-Thinking and LoopCoder-Instruct variants, the first large-scale looped transformer for code, achieving comparable performance to standard dense architectures with more parameters.
Jian Yang, Wei Zhang, Shawn Guo et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.