Vision-language models are increasingly used for driving-scene understanding, yet the semantic relations expressed in their outputs are often difficult to verify against the underlying traffic situation. This paper introduces a deterministic multi-dataset predicate framework that derives driving-scene semantics from me...
Mohamed Chouai, Fazli Faruk Okumus, Stefan Kugele· 0 citations
Three complementary design requirements motivate the present formulation: an explicit force-related constraint, compensation for steady error under persistent loading, and a prediction model that can incorporate trajectory and actuator information.
This work aggregates patch features from all views onto a single equirectangular panoramic canvas, and introduces a spatial pretraining curriculum by procedurally placing patch features of objects at chosen 3D world positions on an otherwise empty canvas, generating on-the-fly supervision spanning a broad range of spat...
Bartłomiej Baranowski, Dave Zhenyu Chen, Matthias Nießner· arXiv.org· 0 citations
Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising alternative to end-to-end vision-language-action systems by combining high-level reasoning with external modules for perception, planning, a...
Haowen Liu, Xirui Li, Shaoxiong Yao et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
End-to-end autonomous driving must balance multimodal maneuver generation against real-time inference constraints. Diffusion planners capture diverse behaviors, but their iterative denoising incurs considerable inference cost in safety-critical deployment, while one-step alternatives still imitate only the single exper...
Yining Xing, Zehong Ke, Zhiyuan Liu et al.· 0 citations
Offline goal-conditioned reinforcement learning requires both long-horizon reachability estimates and local action comparisons. Dual goal representations provide value fields that capture global goal reachability, but they do not directly specify which action should be preferred at a given state. We propose Dual Advant...
Alexey Zemtsov, Maxim Bobrin, Alexander Nikulin et al.· 0 citations
Factorized Contrastive Abstractions for Transferable IRL (ConTraIRL), a framework that enables compositional reward transfer by learning decoupled latent representations of these two factors, is proposed.
This work formalizes this prediction-control interface, separate the scalar trigger from its channel-wise diagnostic log, and establishes that an all-pairs displacement term is redundant within a maximum that already contains the corresponding one-step term.
World Action Models (WAMs) jointly predict future visual observations and actions, but typically execute a fixed number of actions before replanning, regardless of whether the imagined future remains consistent with reality. We formulate adaptive WAM execution as future--reality verification and propose Future Forward...
Rui Wang, Yue Zhang, Canyang Chen et al.· 0 citations
Existing robot video world models are typically trained with low-level objectives such as reconstruction and perceptual similarity, which are poorly aligned with the capabilities that matter most for robot decision making, including instruction following, manipulation success, and physical plausibility. They also suffe...
Fast dexterous grasping while a mobile base remains in motion requires coordinated whole-body control and rapid adaptation to physical contact and this work proposes FastGrasp, a two-stage learning framework that integrates grasp guidance, whole-body control, and tactile feedback.
Heng Tao, Yi-Ming Zhong, Zemin Yang et al.· arXiv.org· 0 citations
Vision-Language-Action models (VLAs) represent a significant frontier in embodied intelligence, aiming to bridge digital knowledge with physical-world interaction. Despite their remarkable performance, foundational VLAs are hindered by the prohibitive computational and data demands inherent to their large-scale archite...
Zhaoshu Yu, Bo Wang, Pengpeng Zeng et al.· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.