A novel architecture is proposed that addresses sample efficiency in VLAs by training a predictive world model on the embedding space of the VLA's vision encoder, and it is hypothesized that these embeddings are action-relevant and usable for future prediction.
Parsa Mastouri Kashani, Jan-Gerrit Habekost, Stefan Wermter· 0 citations
Multi-modal planning is promising for autonomous driving by representing multiple plausible behaviors in ambiguous and long-tail scenarios. Existing methods mainly focus on improving trajectory multi-modality, enhancing trajectory representations, or reshaping the candidate distribution. Nevertheless, we identify a pro...
NavGen is introduced, a text-to-video data generation pipeline that produces diverse vision-language navigation episodes across indoor and outdoor scenes and a style-diversification method that scales up long-tail data that is difficult and costly to collect.
Xi-Jie Huang, Yong-Yang Wan, Cheng-Bin Dong et al.· 1 citation
This work presents a design optimization method that considers both dexterity and anatomy in the design of a bimanual dexterous sheaths robot for performing procedures on cancerous polyps in colon anatomies, achieving a 78% higher RVDSA on average than optimizing for 3D voxel coverage alone.
Tony Qin, Peter Connor, Khoa T. Dang et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
An evaluation protocol is proposed for testing the incremental value of latent access and releasing per-frame failure endpoint labels across two autonomous-driving tasks: online vectorized map generation with LaneSegNet and end-to-end planning with VAD.
N. Advani, Vishwajeet Shivaji Hogale, Saurav Kumar· 0 citations
Closed-Loop contextual Uncertainty rEsolution (Closed-Loop contextual Uncertainty rEsolution), a framework for actively resolving contextual uncertainty given underspecified tasks in natural language, is presented.
Zachary Ravichandran, Jonathan Diller, Fernando Cladera et al.· 0 citations
This paper gives HRI practitioners a standardised basis for evaluating, classifying, and comparing human-robot relationships, and sets out the experimental rigour that each classification demands, and calls on researchers of human-robot relationships to adopt such rigour.
An embodied scene question-answering (QA) interface in which vision never enters the language model is studied, showing the gain survives paraphrase, attributing it to decision-aligned computation rather than answer-string leakage, while novel judgment vocabularies bound its scope.
Albireo is presented, a detector-agnostic, codec-free adaptive inference framework that wraps off-the-shelf detectors and decides when detector invocation can be safely skipped based on scene content and per-object temporal state, requiring no detector modification or retraining.
Amir Taherin, José Cano, Bin Ren et al.· 0 citations
This work proposes a framework that exposes backbone depth V, action expert depth A, and denoising steps $D$ as three jointly configurable compute axes in a VLA, and introduces a KV Cache synthesis mechanism that manages the missing keys and values of the skipped backbone layers, allowing the action expert to exit deep...
Riccardo Andrea Izzo, Rimvydas Rubavicius, Gianluca Bardaro et al.· 0 citations
Results validate the deployment of state-of-the-art anomaly detection on resource-constrained robotics through offline-to-online distillation through a Teacher-Student distillation framework.
Jordan Levy, Nicolas Verstaevel, Vincent Talon et al.· 0 citations
Online reinforcement learning fine-tuning of pretrained flow-matching vision-language-action (VLA) policies promises robots that keep learning after deployment, but continued updates often destroy competence on individual tasks while the aggregate still looks healthy. We study this failure mode, which we call task coll...
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.