Skip to content

Category

robotics

1,156 papers

#artificial intelligence Review Aug 2026

Do World Models Make Better Robots? A Survey of Evaluation Benchmarks for Predictive Embodied Intelligence

An operational taxonomy, a coverage comparison against the eight closest surveys, an evaluation loop that isolates the advantage of prediction, and an actionable protocol of four advantage-aware metrics anchored on named testbeds are contributed.

Gaytri Jena, Kapil Wanaskar, Vinija Jain et al. · 0 citations
#artificial intelligence Preprint Aug 2026

RoboLDA: A Probabilistic Generative Model for Uncovering Embodied Hierarchical Structures in Voxel-based Soft Robots

This work presents RoboLDA, a Bayesian probabilistic model that decomposes VSR morphology generation into a four-level hierarchy:"task-robot-organ-voxel", and is trained via variational inference, which pioneers hierarchical generative modeling of robot morphology.

Jun-Ru Song, Yang Yang, Jing-Dan Shi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HarnessPAI: An Evolving Harness for Physical AI

The results suggest that the frontier of Physical AI depends not only on stronger action models, but also on executable harnesses that integrate perception, task understanding and reasoning, and action execution into a unified, verifiable, and feedback-driven system.

X. Wang, Wen-Hao Wu, Meng-Hao Zhang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

DAWN: Noise-Robust Quadruped Parkour via Depth-Denoising World Models

DAWN is a noise-robust perception framework for legged locomotion which builds noise robustness directly into a world model via two modifications: feeding noisy depth to the encoder while keeping clean depth as the reconstruction target, forcing the model to implicitly denoise its input.

Yo-Han Choi, Min-Jun Kim, Jin-Sung Kim et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Design and Evaluation of LLM Chaining-Based Task Planning for General Purpose Service Robots

This work proposes an LLM chaining architecture that separates instruction classification and action generation into two specialized stages, reducing per-inference prompt length by approximately 45% while improving planning consistency.

Lucas Da Mota Bruno, Jiahao Sim, Y. Hagiwara · 0 citations
#artificial intelligence Preprint Sep 2026

CrossSafe: Towards Cross-Embodiment Latent Safety Filters

This work proposes embodiment-conditioned safety filtering, in which a Hamilton-Jacobi reachability-based value function and its corresponding safety-maximizing policy are shared across robots, and performs Hamilton-Jacobi reachability analysis directly in latent space so that the learned safety concepts can generalize...

Ihab Tabbara, Yuxuan Yang, Hussein Sibai · 0 citations
#artificial intelligence Preprint Sep 2026

Robots That Take Initiative: A Framework for Building and Evaluating Proactive Robots

This work introduces a unified formalism for proactive robot assistance, organize it into three levels, and provides a framework to address the highest level of unprompted proactive assistance, and presents a method, GAP, that instantiates the framework, learning from passive observation to anticipate user goals and ac...

Maithili Patel, Sonia Chernova · 0 citations
#artificial intelligence Preprint Open access Sep 2026

KeyGen: Unsupervised Keypoint based Object-Centric Representations for Category-Level Policy Generalization

Generalization in robotic manipulation requires policies to perform tasks across diverse unseen object instances that vary in shape, size, and pose. However, conventional behavior cloning (BC) methods often overfit to instance-specific geometry and appearance, limiting transfer to novel objects. We introduce KeyGen, a...

Shuxin Cao, Liquan Wang, Masoud Moghani et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Temporal Learning for End-Effector Position Estimation under Aerodynamic Disturbances in Aerial Continuum Manipulation

This paper investigates temporal neural networks for \mbox{end-effector} position \mbox{estimation} of an aerial continuum manipulator (ACM) operating under aerodynamic effects induced by the unmanned aerial vehicle (UAV). An experimental dataset is collected under stationary (\mbox{rotor-off}) and \mbox{free-hovering}...

Niloufar Amiri, Houman Masnavi, Farrokh Janabi-Sharifi · 0 citations
#artificial intelligence Preprint Sep 2026

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

Latent world models are typically trained to predict factual transitions, whereas model predictive control (MPC) must compare alternative actions from the same state. A model can therefore achieve low factual prediction error yet poorly distinguish candidate actions. We introduce AD-WM, an action-discriminative joint-e...

Jia-Bin Qiu, Zi-Xuan Chen, Hong-Ye Cao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Evolve Vision-Language-Action Model into an Agent with On-the-fly Tool-use

ART is a tool-injection framework that tunes any VLA model to leverage off-the-shelf tool modules for low-level vision, high-level affordance, and embodiment enhancement, and ART reduces the complexity of the action solution space through tool-use, which improves generalizability across different tasks but also reduces...

Ding-Ge Yi, Yanzhao Yu, Xi-Li Dai et al. · 1 citation

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.