A high-performing teacher is built that makes navigation evidence selection explicit and compressible, and a compact student is trained by transferring both where to attend and what to do, then further match action distributions during fine-tuning.
Zhihao Chen, Yi-Yuan Ge, Zi-Yang Wang et al.· IEEE transactions on circuit...· 0 citations
CAPEX, an experience-conditioned demonstration collection framework that uses execution experience from previous attempts to adapt how frequently the foundation model must observe, reason, and replan, is introduced and suggests that foundation models can serve as scalable sources of reusable robot experience.
Shivam Aarya, Xi-Jia Zhang, Cheng-Yue Huang et al.· 0 citations
It is found that, given task-specific natural-language guidance, VLMs can estimate actuator-response relationships, resolve ambiguous observations through intervention, and actively create informative visual feedback when signals are sparse, suggesting that pretrained multimodal models can serve as important decision-m...
Ryan Lopez, Sachin Vaidya, Seouing-Ho Choi et al.· 0 citations
The results show that inference-time parameter variation is a reliable, training-free source of path diversity, and that H-Graph hybridization reliably converts this diversity into shorter, higher quality trajectories.
Omer Talmi· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
RoboFoundry is proposed, the first embodied agentic framework that formulates this process of system-as-policy evolution across foundation models as Self-Evolving System-as-Policy, highlighting its potential for fully autonomous embodied agents.
Jing-Song Liang, Shu-Hao Liao, Shi-Zhe Zhang et al.· 0 citations
Ultrasound (US) acquisition depends on the operator's ability to interpret anatomy and anticipate how the view will change with probe motion. Many robotic US navigation methods select actions without explicitly predicting these anatomical changes. We propose SonoGraph-WM, an action- and goal-conditioned world model for...
Xue-Song Li, Shuai Chen, Feng Li et al.· 0 citations
RECAST is a robot navigation framework that combines the reasoning of a VLM with the spatial grounding of vision foundation models to build an Actionable Cost map and improves success over the strongest prior method.
Incheol Cho, Jintae Park, Jinkyu Kim et al.· 0 citations
Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a short temporal window. Equipping such policies with long-term memory remains challenging: existing approaches either feed the backbone multi-frame observation windows, which substant...
Xin-Yu Zhao, Yixiang Shan, Tao Yang et al.· 0 citations
Code-as-Policy agents accomplish long-horizon embodied tasks by generating and executing code, yet continually improving them with teachers that are stronger but not globally reliable remains a key challenge. Existing distillation methods typically treat the teacher's complete behavior as the supervision target and thu...
Jiawei Zhang, Xiangrong Zhang, Rui Song et al.· 0 citations
Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into reliable physical execution remains difficult in spatially constrained transitions. Reaching a staircase, doorway, or narrow passage does not ensure traversal; the agent must identify...
Xue-Kang Yang, Lu-Yao Chen, Shuang Luo et al.· 0 citations
DS-VLA is introduced, a dendritic-inspired action architecture that incorporates dendritic spiking dynamics into VLA control to enable modularized feature processing and temporal information integration and demonstrate that integrating brain-inspired computational mechanisms offers a promising architectural prior for r...
When several surfaces are visible along a ray, recovering visible 3D structure from one image requires jointly estimating their presence and metric depth. Depth Any Seen represents these surfaces as image-conditioned multi-Bernoulli depth sets, whose components each contribute one depth or remain absent. Its auxiliary-...
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.