Pretrained video diffusion models provide powerful spatiotemporal generative priors, making them a natural foundation for robotic world models. While recent world-action models jointly optimize future videos and actions, they predominantly treat video generation as an auxiliary representation for policy learning. Conse...
Zhaoyang Yang, Yurun Jin, Lizhe Qi et al.· 0 citations
Monocular 3D scene reconstruction has recently seen significant progress. Powered by the modern neural architectures and large-scale data, recent methods achieve high performance in depth estimation from a single image. Meanwhile, reconstructing and decomposing common scenes into individual 3D objects remains a hard ch...
Junaid Ahmed Ansari, Ran Ding, Fabio Pizzati et al.· 0 citations
We develop a decentralized Perception-Action-Communication (PAC) system for multi-robot teams that enables them to collaborate in large scale, outdoor environments. Our system natively supports deployments at any scale by leveraging a graph neural network (GNN) to diffuse information hop-by-hop across the fleet's netwo...
Social navigation typically assumes a specified goal and focuses on reaching it while respecting social conventions, whereas robot group joining requires predicting where to join based on the group's real-time activity and formation. This is a highly semantic task, yet an important capability for applications such as r...
Zilin Fang, Zishuo Wang, Gim Hee Lee et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Latent world models that integrate a flow in a frozen self supervised latent space train stably and cheaply, yet silently lose the property manipulation depends on most: motion. The pretrained flow never moves the manipulated object; retraining it with latent-only losses only trades stillness for teleport-like motion....
Xiwen Chen, Rigaudiere Z. Li, Zhiruo Zhou et al.· 0 citations
Vision-Language-Action models provide a strong foundation for general-purpose robot control, yet a vast majority of policies do not preserve and leverage episode-level information beyond the current observation. This limitation is consequential in history-dependent manipulation tasks that depend on information availabl...
Tej Deep Pala, Navonil Majumder, Bryce Goh et al.· 0 citations
Large Language Models (LLMs) introduce an exciting new paradigm for planning and navigation in robotics, but fail on even simple multi-robot tasks as team sizes grow. We propose COMPASS, a scalable, decentralized multi-robot architecture for controlling large collectives of agentic robots with reasoning space feedback...
Frederic Vatnsdal, Roshan Gopal, Romina García Camargo et al.· 0 citations
This study presents a centralized safety aware M2M framework for cooperative goal-directed navigation in a heterogeneous mobile robot system composed of a vision-capable robot and a cameraless robotic vehicle that can approve safe motion, trigger conservative replanning or holding behavior, and preserve a strict separa...
Mohamed Dwedar, Ahmad Hafez, Alexander Jesser et al.· IEEE Access· 0 citations
InfiNoVA is introduced, a data-augmentation framework that converts synchronized multi-camera demonstrations into a dense distribution of geometrically consistent training views that improves frame-level fidelity and temporal consistency while reducing task-critical hallucinations observed in generative novel-view synt...
Sai Puneeth Reddy Gottam, Elmar Rueckert, Vedant Dave· 0 citations
Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal interfaces, asynchronous multi-...
Ji-Song Cai, Yao Mu, Gan-Lin Yang et al.· 0 citations
Behaviora is a preliminary conceptual architecture for representing agent and robot behavior, external and internal alike, in an addressable form. A behaving robot or agent performs a Behavior Episode composed of episode components, which can be derived from behavior taxonomies (BTax) and assigned persistent identifier...
Long-term autonomy in human-populated environments requires anticipating whether and how people will move at times a robot has not yet observed. Existing representations of pedestrian motion face a tradeoff: they either forecast future activity, reducing each location to a scalar rate, or model the full directional dis...
Iacopo Catalano, Julio A. Placed, Javier Civera et al.· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.