Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale errors undo all prior progress. Online reinforcement learning (RL) can optimize exactly these actions, but free exploration is far too costly on real robots, which makes hum...
Wei-Hui Zhao, Xiao Yan, Zu-Nian Wan et al.· 0 citations
Automotive spinning FMCW radar provides dense, $360^\circ$ sensing and remains reliable under poor illumination and adverse weather, making it well-suited to autonomous navigation. Place recognition uses these observations to identify previously visited locations for re-localization and long-term navigation. However, h...
Saimunur Rahman, Sagun Shrestha, Abdelwahed Khamis et al.· 0 citations
Visual Place Recognition (VPR) in natural environments remains challenging due to repetitive vegetation, sparse distinctive landmarks, and substantial appearance and viewpoint variation across traversals. While visual observations of the same place can change considerably, their underlying spatial structure is often mo...
Walter Nedov, Saimunur Rahman, Kavindie Katuwandeniya et al.· 0 citations
It is found that fine-tuned imitation learning policies converge to similar action predictions across priors, despite their fine-tuned encoder embeddings diverging substantially from the pretrained embeddings and each other.
Chen Xu, Rishi Shah, H. Kress-Gazit et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Navigation in off-road conditions is challenging due to the lack of structure. There is no fixed vocabulary for what is traversable. The traversability depends on both the environment and the embodiment's dynamics. Neither of these two variables can be hand-labeled at scale. Thus, traversability has to be learned by th...
W. Bonilla, M. Boisvert, David-Alexandre Poissant et al.· 0 citations
SlackDrive is proposed, a pre-inference compute allocator that reuses realized latency to select the compute budget of each control step before model execution, complementing existing profiling and resource scheduling while preserving the driving backbone and its compute actuator.
Xiao-Huan Pei, Heng-Guang Zhou, Yuan-Hao Ban et al.· 0 citations
Talk2Escape is introduced, a proactive and model-agnostic dialogue intervention framework that reframes navigation as a closed-loop interactive process and proves that proactive dialogue drastically improves navigation robustness in physical environments.
Ze-Rui Li, Si-Hao Lin, Yan-Yan Shao et al.· 0 citations
Electrovibration-based tactile feedback is demonstrated to be a viable and effective modality for robot teleoperation, improving operator responsiveness and sense of presence in contact-rich manipulation tasks, with direct applicability to safety-critical domains such as nuclear maintenance.
Alperen Kenan, Juan Jose Garcia Cardenas, Adriana Tapus et al.· 0 citations
This work proposes DreamAvoid, a critical-phase test-time dreaming framework that enables VLA models to anticipate and avoid failures, and introduces an autonomous boundary learning paradigm to refine the system's understanding of the subtle boundary between success and failure.
Across towel folding, book insertion, and Hanoi ring placement, LIFT learns faster and reaches higher performance than vision-only post-training, while ablations show that reactive force memory and online corrective data are both important for robust contact-rich manipulation.
Yi Wang, Wen-Di Chen, Zi-Mo Wen et al.· arXiv.org· 4 citations
Existing texture datasets for tactile sensing primarily consist of sensor readings from a specific sensor interacting with available surfaces/objects rather than describing the textures themselves, limiting fair comparison between tactile sensors and hindering reproducible research. In this work, we introduce a 3D-prin...
Dexter R. Shepherd, Nicolas Herzig, Phil Husbands et al.· 0 citations
Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over spatiotemporal structure suc...
Maijunxian Wang, Ruisi Wang, Juyi Lin et al.· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.