This work introduces Anchored Planning, which retrieves a recorded segment whose start and end resemble the current and goal observations, then aims at an observation shortly after its start, and without additional training, planning toward observed targets outperforms the LeWM planner on every task in the authors' lon...
Xvyuan Liu, Jian-Jie Fang, Chen Gao et al.· 0 citations
Preliminary results indicate that a world-action model (WAM) post-trained from OmniDreams achieves strong performance on the Physical AI Autonomous Vehicles NuRec dataset, surpassing the VLA-based Alpamayo 1.5 research policy model while using only 1/5 the total parameters.
Large language models (LLMs) are increasingly used as proposal generators for evolutionary robot design, yet most loops remain memoryless: simulator results shape the next population but are not preserved as reusable design knowledge. We present Auto-Robotist, a self-evolving LLM agent that distills morphology-search t...
Yunfei Wang, Xiaohao Xu, Yang Li et al.· 0 citations
This work proposes a neuro-symbolic architecture that integrates symbolic planning, reinforcement learning, and a large language model (LLM) to learn how to handle novel objects and outperforms the state-of-the-art methods in operator discovery as well as operator learning in continuous robotic domains.
Hongxuan Lu, Pierrick Lorang, Timothy R. Duggan et al.· arXiv.org· 2 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This study systematically dissects design choices along three dimensions: foundational components, perception essentials, and action modeling perspectives and distill 12 key findings that together form a practical recipe for building strong VLA models, resulting in a simple yet effective model, VLANeXt.
Xiao-Ming Wu, Kang Liao, Yi-Hang Luo et al.· 12 citations· ⚡3
This work introduces Robot Agentic Programming from Demonstrations (RAPID), which automatically generates, verifies, and refines robot programs, given a single visual human demonstration, using an object-centric relational program representation.
Yu-Yao Liu, Jia-Yuan Mao, David Hsu et al.· 0 citations
Rolling-WAM is presented, a formulation that distributes joint denoising across successive replanning cycles and delivers a 4.5x steady-state replanning speedup over standard joint WAMs.
Ying Zhou, Jun-Jie Ye, Yi-Qi Zhao et al.· 0 citations
This work finds that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success.
Matteo Merler, Bo-Wen Li, Josh Roy et al.· 1 citation
TrackEverything is the first 3D tracker capable of tracking all visible points across videos exceeding 1000 frames within 40 GB of GPU memory, and 3D WAFT is proposed, replacing memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud.
Ayush Jain, Sreeharsha Paruchuri, Ishita Gupta et al.· 0 citations
We present Underwater C$^{3}$-JEPA (cross-view, control-conditioned, context-extended), an object-centric multi-view predictive world model for near-field heavy-load underwater ROV salvage. Without contact sensors, it predicts in latent space how the task-object state evolves through contact interaction and under the h...
Yuncong Yang, Jin-Long Li, Yu-Long Xue et al.· 0 citations
World Action Agent is presented, a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.
Ye-Hang Zhang, Hao-Jian Huang, Yi-Fan Chang et al.· 0 citations
MorphIK is a flow-matching model that solves inverse kinematics for revolute-joint-based kinematic chains it has never seen during training, and allows learning and generalizing neural inverse kinematics for a multitude of known and unknown robots.
Lennart Clasmeier, Jan-Gerrit Habekost, C. Weber et al.· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.