Skip to content

Category

robotics

1,156 papers

#machine learning Preprint Sep 2026

Grounding Vision-Language Models in Driving Semantics: A Multi-Dataset Predicate Framework for Explainable Reasoning

Vision-language models are increasingly used for driving-scene understanding, yet the semantic relations expressed in their outputs are often difficult to verify against the underlying traffic situation. This paper introduces a deterministic multi-dataset predicate framework that derives driving-scene semantics from me...

Mohamed Chouai, Fazli Faruk Okumus, Stefan Kugele · 0 citations

OneCanvas: 3D Scene Understanding via Panoramic Reprojection

This work aggregates patch features from all views onto a single equirectangular panoramic canvas, and introduces a spatial pretraining curriculum by procedurally placing patch features of objects at chosen 3D world positions on an otherwise empty canvas, generating on-the-fly supervision spanning a broad range of spat...

Bartłomiej Baranowski, Dave Zhenyu Chen, Matthias Nießner · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Guava: Distilling Frontier VLM Agents into a Compact Model with a Manipulation Harness

Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising alternative to end-to-end vision-language-action systems by combining high-level reasoning with external modules for perception, planning, a...

Haowen Liu, Xirui Li, Shaoxiong Yao et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Vault: One-Step Latent Generation with Positive-Anchored Rewards for Autonomous Driving

End-to-end autonomous driving must balance multimodal maneuver generation against real-time inference constraints. Diffusion planners capture diverse behaviors, but their iterative denoising incurs considerable inference cost in safety-critical deployment, while one-step alternatives still imitate only the single exper...

Yining Xing, Zehong Ke, Zhiyuan Liu et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Dual Advantage Fields

Offline goal-conditioned reinforcement learning requires both long-horizon reachability estimates and local action comparisons. Dual goal representations provide value fields that capture global goal reachability, but they do not directly specify which action should be preferred at a given state. We propose Dual Advant...

Alexey Zemtsov, Maxim Bobrin, Alexander Nikulin et al. · 0 citations

Can Predicted Dynamics Exist in the Physical World?

This work formalizes this prediction-control interface, separate the scalar trigger from its channel-wise diagnostic log, and establishes that an all-pairs displacement term is redundant within a maximum that already contains the corresponding one-step term.

Barak Or · 0 citations
#artificial intelligence Preprint Open access Sep 2026

When to Trust Imagination: Adaptive Action Execution for World Action Models

World Action Models (WAMs) jointly predict future visual observations and actions, but typically execute a fixed number of actions before replanning, regardless of whether the imagined future remains consistent with reality. We formulate adaptive WAM execution as future--reality verification and propose Future Forward...

Rui Wang, Yue Zhang, Canyang Chen et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models

Existing robot video world models are typically trained with low-level objectives such as reconstruction and perceptual similarity, which are poorly aligned with the capabilities that matter most for robot decision making, including instruction following, manipulation success, and physical plausibility. They also suffe...

Hao Wu, Yuqi Li, Yuan Gao et al. · 0 citations

FastGrasp: Learning-based Whole-body Control method for Fast Dexterous Grasping with Mobile Manipulators

Fast dexterous grasping while a mobile base remains in motion requires coordinated whole-body control and rapid adaptation to physical contact and this work proposes FastGrasp, a two-stage learning framework that integrates grasp guidance, whole-body control, and tactile feedback.

Heng Tao, Yi-Ming Zhong, Zemin Yang et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

A Survey on Efficient Vision-Language-Action Models

Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale archite...

Zhaoshu Yu, Bo Wang, Pengpeng Zeng et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.