This work presents Uranus, a data-driven robot simulator built around a joint-trajectory-conditioned autoregressive diffusion model, providing a unified interface for synchronized multi-view generation across diverse robot embodiments and camera configurations.
Wen-Kang Qin, Yu-Kun Zhou, Noah Shen et al.· 0 citations
Humanoid robots are increasingly being popular and developed for human-centered applications, yet their ability to provide intelligent conversations and natural interactive knowledge assistance remains constrained by traditional rule-based dialogue systems, pre-defined responses and limited knowledge repositories. Larg...
Human demonstrations offer a scalable way to collect manipulation data, but their contacts may be unstable or infeasible when transferred to a robot hand. Collecting demonstrations directly on the target robot avoids this mismatch but substantially increases the cost of data collection. To address this trade-off, we pr...
Sheng-Cheng Luo, Xiao Cheng, Hong Ying et al.· 1 citation
Autonomous parking in nonconvex and narrow environments remains challenging. Although optimal-control methods can explicitly enforce vehicle dynamics and collision constraints, nonconvexity compromises solver robustness and can cause failures. Large language models (LLMs) exhibit strong semantic reasoning capabilities,...
Recent 4D LiDAR language models aim to reason about objects and their evolving spatial relationships. Yet, in our evaluation, always selecting the same option nearly matches the multiple-choice accuracy of two B4DL-derived configurations. We introduce LiDAR-Hallu, a geometry-referenced benchmark and diagnostic protocol...
Runyi Yang, Murat Akkoyun, Di Wen et al.· 0 citations
Low-bit vision-language-action inference must reduce observation-to-action latency while preserving robot behavior. We present FoldQuantVLA, a post-training quantization framework that carries a consistent activation representation through calibration, weight rounding, and native integer execution. It combines channel...
Hung T. Ho, Khanh-Binh Nguyen, Quan Nguyen et al.· 0 citations
Tactile-JEPA is an efficient self-supervised pre-training method that uses the spatial arrangement of tactile sensors to learn topology-aware representations, and is trained to predict the embeddings of masked sensing elements from the unmasked remainder, using the sensor connectivity graph to guide spatial masking.
E. Kovtun, M. Konovalov, Andrey Sakhovskiy et al.· 0 citations
This work presents vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization, and introduces IMPACT, an ACT-based policy with cached text representations and language-modulated visual features.
Khanh Duy Nguyen, Hoang M. Truong, A. T. Le· 0 citations
The ActiveArena-Sim, an active-perception simulator with controllable viewpoints and large-scale workspaces as the foundation, is introduced, and ActiveArena-Bench is proposed, which comprises 35 tasks across 5 fine-grained categories, covering visual exploration and interactive information acquisition.
Yi-Bo Li, En-Shen Zhou, Rui Chen et al.· 1 citation
Simulation-based evaluation provides a scalable and repeatable alternative to real-world evaluation of vision-language-action (VLA) policies. However, reconstruction errors can cause simulated policy performance to diverge from real-world performance, motivating the need to assess reconstructed environments for downstr...
Xinyi Wang, Heng Hao, Wenjun Hu et al.· 0 citations
The spiking neural network-based Proximal Policy Optimization algorithm integrates the use of spike-based actor-critic reinforcement learning with the Proximal Policy Optimization algorithm and uses the stochastic Gaussian policy in the autonomous navigation of unmanned aerial vehicles.
F. Walugembe, Maciej Wielgosz, Tomaž Goričan et al.· 0 citations
This work introduces RiverVLN, to its knowledge the first benchmark designed for long-horizon USV VLN under continuous riverine motion, and PGT-NAV, a phase-grounded temporal navigation framework for USVs.
Jie-Ling Wu, Yue-Hao Huang, Jia-Jun Lv et al.· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.