Skip to content

Category

robotics

1,156 papers

#artificial intelligence Preprint Open access Oct 2026

Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory

Vision-Language-Action (VLA) models are increasingly deployed across diverse tasks, yet they remain black boxes whose physical interactions can cause irreversible harm, making generalizable and interpretable failure detection essential. We observe that successful and failed rollouts carry systematically different infor...

Jinghan Yang, Yunchao Zhang, Wang Yuan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ATM: Why Latent World Models Can Fail to Plan

Latent world models can achieve accurate latent prediction yet still differ substantially in downstream planning performance. We argue that a key source of this discrepancy lies in the structure of action-induced latent transitions. We formalize action-identifiability through Bayes inverse risk, characterizing how much...

Jiaheng Chen, Tinghe Zhang, Yucheng Xiao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Understanding Multimodality in Generative Behavioral Cloning

Behavioral cloning becomes challenging when the same observation admits several valid actions. We study how generative behavioral-cloning policies represent such multimodal expert behavior and identify different bottlenecks across model parameterizations. For latent-variable policies, preserving demonstrated modes requ...

Lorenzo Mazza, Massimiliano Datres, Ariel Rodriguez et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DexHoldem: An Agentic Robotics Benchmark for Dexterous Manipulation in Texas Hold'em

Evaluating embodied systems with real dexterous hardware requires more than isolated motor-skill tests: an agent must perceive a changing scene (e.g. a tabletop), choose a context-appropriate action, execute it with a dexterous hand, and leave the scene usable for later decisions. We introduce DexHoldem, a comprehensiv...

Feng Chen, Tianzhe Chu, Li Sun et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Bimanual Robot Manipulation via Multi-Agent In-Context Learning

Large Language Models (LLMs) have emerged as powerful reasoning engines for embodied control. In particular, In-Context Learning (ICL) enables off-the-shelf, text-only LLMs to predict robot actions without any task-specific training while preserving their generalization capabilities. Applying ICL to bimanual manipulati...

Alessio Palma, Indro Spinelli, Vignesh Prasad et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Grounded World Model: Latent Planning with Language Goals

World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks. We present the Grounded World Model (GWM), a latent world model that enables zero-shot planning in the real world from language goals alone. Given a candidate action sequence and the current ob...

Quanyi Li, Lan Feng, Haonan Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Language-Conditioned World Modeling for Visual Navigation

Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, language-conditioned visual navigation (LCVN), in which an embodied agent must follow a natural language instruction given only an initial egocentric observation. Without access to goal...

Yifei Dong, Fengyi Wu, Yilong Dai et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ExpertGen: Scalable Sim-to-Real Expert Policy Learning from Imperfect Behavior Priors

Learning generalizable and robust behavior cloning policies requires large volumes of high-quality robotics data. While human demonstrations (e.g., through teleoperation) serve as the standard source for expert behaviors, acquiring such data at scale in the real world is prohibitively expensive. This paper introduces E...

Zifan Xu, Ran Gong, Maria Vittoria Minniti et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Pretrained robot policies provide useful action priors, but long-horizon manipulation still requires coordination between semantic reasoning and physical execution. Semantic reasoning operates at a coarser timescale than physical interaction, while episode-level failures provide limited guidance on which system compone...

Hao-Yuan Deng, Jie-Bin Liu, Teng-Xiao Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

We present PrefPI (Preference-Guided Policy Iteration), an iterative framework for steering pretrained generative robot policies using only relative preferences over self-generated trajectories. Unlike prior preference-learning methods that primarily sharpen modes already represented by the policy, we study steering be...

Seungeun Rho, Wontaek Kim, Danfei Xu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Game-Guided Skill Discovery through Self-Play for Playable Agent Control

We present Game-Guided Skill Discovery (GGSD), a framework that uses self-play in games to discover motor skills that are directly playable by humans. Playable skills provide a compact abstraction for controlling embodied agents through a small set of learned behaviors rather than low-level actions. To be effective, th...

Seungeun Rho, Jeonghwan Kim, Xue Bin Peng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Tactile Curiosity Drives Robot Interaction

Mastering robot manipulation skills via reinforcement learning (RL) remains largely sample-inefficient. The most common RL algorithms rely on random action sampling to discover new strategies, resulting in agents that allocate most of their training budget to motions in free space, away from the contacts from which man...

Klemens Iten, Alexander Proshkin, Bhavya Sukhija et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.