Skip to content

Category

robotics

1,156 papers

#artificial intelligence Open access Sep 2026

Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation

A high-performing teacher is built that makes navigation evidence selection explicit and compressible, and a compact student is trained by transferring both where to attend and what to do, then further match action distributions during fine-tuning.

Zhihao Chen, Yi-Yuan Ge, Zi-Yang Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CAPEX: Efficiently Distilling Foundation Model Behavior into Deployable Robot Policies through Experience-Adaptive Reasoning

CAPEX, an experience-conditioned demonstration collection framework that uses execution experience from previous attempts to adapt how frequently the foundation model must observe, reason, and replan, is introduced and suggests that foundation models can serve as scalable sources of reusable robot experience.

Shivam Aarya, Xi-Jia Zhang, Cheng-Yue Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Vision-Language Agents for Active Perception in Optics Laboratories

It is found that, given task-specific natural-language guidance, VLMs can estimate actuator-response relationships, resolve ambiguous observations through intervention, and actively create informative visual feedback when signals are sparse, suggesting that pretrained multimodal models can serve as important decision-m...

Ryan Lopez, Sachin Vaidya, Seouing-Ho Choi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Optimizing H-Graph Hybridization for Diffusion-Guided RRT

The results show that inference-time parameter variation is a reliable, training-free source of path diversity, and that H-Graph hybridization reliably converts this diversity into shorter, higher quality trajectories.

Omer Talmi · 0 citations
#artificial intelligence Preprint Sep 2026

RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents

RoboFoundry is proposed, the first embodied agentic framework that formulates this process of system-as-policy evolution across foundation models as Self-Evolving System-as-Policy, highlighting its potential for fully autonomous embodied agents.

Jing-Song Liang, Shu-Hao Liao, Shi-Zhe Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Scanning While Imagining: A Scene-Graph World Model for Robotic Ultrasound Navigation

Ultrasound (US) acquisition depends on the operator's ability to interpret anatomy and anticipate how the view will change with probe motion. Many robotic US navigation methods select actions without explicitly predicting these anatomical changes. We propose SonoGraph-WM, an action- and goal-conditioned world model for...

Xue-Song Li, Shuai Chen, Feng Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DRAM: Delta-rule Recurrent Associative Memory for Robot Manipulation Policies

Robotic manipulation is inherently history-dependent, yet most pretrained robotic policies condition on only the current observation or a short temporal window. Equipping such policies with long-term memory remains challenging: existing approaches either feed the backbone multi-frame observation windows, which substant...

Xin-Yu Zhao, Yixiang Shan, Tao Yang et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

RE-0: Verified Recursive Improvement of Embodied Code-as-Policy Agents through Local On-Policy Distillation

Code-as-Policy agents accomplish long-horizon embodied tasks by generating and executing code, yet continually improving them with teachers that are stronger but not globally reliable remains a key challenge. Existing distillation methods typically treat the teacher's complete behavior as the supervision target and thu...

Jiawei Zhang, Xiangrong Zhang, Rui Song et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Affordance-Conditioned Decision Making: Bridging the Semantic-Spatial Gap in Zero-Shot Cross-Floor Vision-and-Language Navigation

Vision-and-language navigation increasingly relies on general-purpose semantic planners, yet translating correct high-level intent into reliable physical execution remains difficult in spatially constrained transitions. Reaching a staircase, doorway, or narrow passage does not ensure traversal; the agent must identify...

Xue-Kang Yang, Lu-Yao Chen, Shuang Luo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DS-VLA: A Dendritic-inspired Vision-Language-Action Model for Robust Action Control

DS-VLA is introduced, a dendritic-inspired action architecture that incorporates dendritic spiking dynamics into VLA control to enable modularized feature processing and temporal information integration and demonstrate that integrating brain-inspired computational mechanisms offers a promising architectural prior for r...

Yaxing Lyu, Jing-Yi Li, M. Xu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Depth Any Seen: Which Surfaces and How Far?

When several surfaces are visible along a ray, recovering visible 3D structure from one image requires jointly estimating their presence and metric depth. Depth Any Seen represents these surfaces as image-conditioned multi-Bernoulli depth sets, whose components each contribute one depth or remain absent. Its auxiliary-...

Xiao-Hao Xu, Xiao-Nan Huang · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.