Skip to content

Category

robotics

1,156 papers

#artificial intelligence Preprint Open access Oct 2026

BatSLAM 2.0: Sequence-Verified Sonar Place Recognition in a Robust Pose Graph

Echolocating bats can navigate dark and cluttered spaces using echolocation. Over a decade ago, BatSLAM showed that a robot with a biomimetic binaural sonar can build a topological map of the environment, by recognizing places from the received acoustic signals. Sonar place recognition, however, is ambiguous by nature:...

Jan Steckel · 0 citations
#artificial intelligence Preprint Sep 2026

TACTIC: Temporal and Context-Aware LLM Tactical Planning for Roadside LiDAR Attacks

Physical LiDAR attacks are often evaluated using fixed primitives and manually selected parameters, despite their strong dependence on surrounding traffic. We present TACTIC, a scene-aware framework that uses a multimodal large language model (MLLM) to coordinate state-adaptive roadside LiDAR attacks. Under a gray-box...

Yi-Ming Gao, Shao-Cheng Luo · 0 citations
#artificial intelligence Preprint Sep 2026

Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models

Vision-language-action (VLA) models generalize broadly across robotic manipulation tasks, but complex environments require balancing task success with unintended contact. Runtime shields can correct individual actions, but they leave the underlying policy unchanged, so repeated disagreements may create a persistent pol...

M.-Y. Cui, Zhe-Yuan Liu, Yi-Han Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DiffWAM: A Fast and Efficient Navigation World Action Model

Pretrained video foundation models encode rich semantic and spatiotemporal priors for embodied navigation, yet converting these priors into UAV motion typically requires expensive future-video synthesis and geometric reconstruction. We investigate whether the motion implicit in future visual prediction can instead be r...

Morui Zhu, Yu-Ze Wu, Xi-Jie Huang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

RoboCoach: World Models as Active Coaches for Compositional Robot Skills

Long-horizon robot manipulation reuses skills across many task compositions, but improving these compositions with additional end-to-end demonstrations is costly. A practical self-improving system must decide both what to teach next and where to apply that supervision. We present ROBOCOACH, a world-model-guided coachin...

Jia-Jun Liu, Yi-Fan Chen, Yi-Chao Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generati...

Qi-Ze Yu, Lian-Rui Fan, Bo-Yu Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply...

Qize Yu, Lianrui Fan, Bowen Ping et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization

3D visuomotor policies provide a strong foundation for spatially precise manipulation, yet current text-to-3D policies struggle to follow unseen fine-grained behavioral specifications beyond those covered by demonstrations. We study this challenge as unseen specification generalization, where language specifies behavio...

Xinhao Yang, Wenhao Wu, Ning Lv et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ECHO-G: Embodied Co-speech Humanoid mOtion Generation

Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-alig...

Yi-Zhao Li, Pu-Sen Gao, Ming Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HiWE: Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning

Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-re...

Guo-Qing Ma, Ming-Qi Yuan, Chen Gao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Scale and Selection: What Makes Automatic Harness Evolution Work for Visual-Interface Robot Agents

When an off-the-shelf coding agent is used directly as a robot policy, observing a browser-based 3D interface through screenshots and acting by posing a virtual target gripper through a few tools, the agent's harness, its prompts, tools, and control rules, largely determines success, and until now it has been written b...

Zhijie Wei, Ferris Tan, Jinghui Wang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Blackout vs. Freeze: Analyzing Physical Failure Modes of VLAs under Camera Faults

Unreliable visual inputs can harm task performance and cause potential physical safety risks for vision-language-action (VLA) models. We analyze how $\pi 0.5$ and GR00T models act under input faults such as image blackouts and freezing. We find that blackout and freezing produce distinct physical failure modes even whe...

Heejae Suh, Jongwook Han, Zahra Gholami et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.