Skip to content

Category

robotics

1,156 papers

#artificial intelligence Preprint Oct 2026

Video2World: Benchmarking Coding Agents for Interactive World Modeling from Embodied Videos

Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily on manual environment construction and calibration. We study whether frontier foundation models and coding agents can automate this process end to end. We formulate \emph{auto...

Jin-Zhou Tang, Z. Zhang, Jing Yang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Human Behavior-Informed Crash Scenario Generation with Real-World Crash Priors for Autonomous Vehicle Safety Evaluation

Reliable safety evaluation of autonomous vehicles (AVs) is essential to improving road safety, yet it depends critically on realistic simulation of rare crashes. Existing crash scenario generation methods can increase collision occurrence, but often fail to realistically reproduce how crashes evolve before impact or th...

Mingxing Peng, Xusen Guo, Long Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Bird's-Eye View of Iterative Reward Design

Designing effective reward functions in RL typically requires substantial expertise and trial and error. Recent work automates this process with LLM-based systems that generate and iteratively improve reward code using policy feedback. However, these methods are often hard to compare because they differ in implementati...

Logan Mondal Bhamidipaty, Lauren Robson, Linda Petrini et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Attention-Based Surface Representation Learning for Robot State Prediction and Open-Ended Surface Classification

For ground robots operating in outdoor environments, understanding the properties of the underlying terrain is essential for ensuring reliable operation. In most perception-based studies, this problem is formulated as categorical classification with a fixed number of classes defined during training. We propose an appro...

Oleg Kushnarev, Alexander Belyaev · 0 citations
#artificial intelligence Preprint Open access Oct 2026

CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model

Skeletal motion is stored as every joint's transform at every frame, yet most of it is implied by the body rather than by what the motion is about. Compression is one way to ask what a motion must still say once the body is known, and a production codec must answer it for any skeleton with a stated error bound. Our ear...

Mingyi Shi, Huancheng Lin, Xuelin Chen et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models

Recent advances in robot learning have enabled generalist control policies capable of completing a wide range of tasks. However, their performance degrades when deployed in unseen environments, making it critical to detect failures and teach recovery behaviors. Existing runtime monitoring methods often require task- an...

Ryan Li, Yigit Korkmaz, Erdem Biyik · 0 citations
#artificial intelligence Preprint Oct 2026

GOTT: Object-centric Dexterous Manipulation with a Reusable Cross-Embodiment Primitive

Foundation models and large-scale human data provide rich sources of manipulation intent, but translating this intent into multi-fingered robot behavior remains difficult. Dexterous hands still lack a reusable low-level primitive that reliably establishes contact across tasks and embodiments. We propose GOTT, a reach-a...

Yu-Lin Liu, Lai Wei, Yen-Jen Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

WAMJET: A Harness for World Action Model Acceleration

World Action Models (WAMs) leverage pretrained video foundation models for robot manipulation, but their large backbones and video-action co-prediction are expensive. Although existing acceleration techniques offer many ways to reduce this cost, selecting and composing them requires substantial engineering for each mod...

Le Chen, Lixin Liu, Jan Schneider et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies

A safe action is not necessarily a viable one. Under a frozen vision-language-action (VLA) policy, an action can be likely and locally admissible yet leave no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likelihood ranks the current action, whereas feasibility depends on...

Tu Nguyen, Matthieu Zimmer, Vu Anh Vu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

PermVLA: Factorization Order as a Regularizer for VLA Learning

Vision-language-action (VLA) policies commonly learn action chunks through a fixed left-to-right (LTR) factorization, although the same expert trajectory distribution admits many valid chain-rule factorizations. We identify factorization order as an overlooked regularization choice and introduce causally anchored permu...

Yanqiao Chen, Yuhan Rui, Dongsheng Hou et al. · 0 citations
#artificial intelligence Preprint Oct 2026

RAGrasp: Geometry-Semantic Template Retrieval and Grasp Transfer

We present RAGrasp, a retrieval-augmented pipeline for planar parallel-jaw grasping from a compact set of locally collected, grasp-annotated RGB-D (color and depth) templates. Unlike task-specific predictors trained primarily on large public or synthetic grasp datasets, RAGrasp requires no end-to-end retraining for a n...

Shen-Zhe Zhu, Cheng-Xiao He, Jan Harder · 0 citations
#artificial intelligence Preprint Oct 2026

Towards Safer Autonomous Driving in an Open World: A Dual-Process Approach

Before autonomous driving systems can be deployed on public roads, it is vital that these systems comply with safety standards, traffic rules, and social norms. Although neural networks trained on large amounts of driving data perform well in routine driving tasks, these models often struggle in novel situations that a...

Simon Janssen, Michiel Braat, C. van der Ploeg et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.