Skip to content

Category

robotics

1,156 papers

#artificial intelligence Preprint Sep 2026

Robot-GST: geometry-aware spatial-temporal robot policy representation and evaluation

Robotic-GST is presented, a geometry-aware spatio-temporal behaviour representation and evaluation framework that constructs a Gaussian-SAM robotic environment for real-to-sim policy verification and improves the reliability of real-world manipulation deployment.

Si-Chao Liu, Ze-Kun Wang, Li-Xuan Tang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Achieve What You Imagined: Learning to Align Actions with Visual Plans

World-action models can jointly predict future visual observations and robot actions. However, discrepancies may exist between their visual predictions and the consequences implied by generated actions. We observe that WAMs can often generate visually plausible task-completion outcomes before producing action sequences...

Yu-Heng Qiao, Zi-Ran Wei, Xiao-Hang Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation

CodeActionBench is introduced, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy and provides a controlled testbed for measuring how general-purpose models translate their capabilities into manipulation behavior and for examining typical failure scenarios in that process...

Yiheng Lyu, Xueying Jiang, Wen-Hao Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Demonstration-Free Success-Probability Reward Learning for Generalist Robot Policies

This work introduces eVTA, which learns success probabilities from mixed-quality policy rollouts through temporal-difference-style bootstrapping, without expert demonstrations or intermediate annotations, and introduces RL with Evolving Rewards (RLER), a closed-loop framework that adapts eVTA using newly collected roll...

Duo Wu, Hai-Feng Wang, Rongwei Lu et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

AquaWAM: A Dynamics-aware World Action Model for Underwater Embodied Agents

World Action Models (WAMs) are becoming increasingly important and useful for embodied intelligence, as they enable robots to anticipate the consequences of candidate actions before interacting with the physical environment. However, underwater robots are usually subject to passive dynamics, such as inertia, buoyancy,...

Cunhao Zhu, Yifeng Wang, Dongliang Xu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Scope-WM: Scoped Computation for Efficient Visual World Models

Visual world models enable robotic planning by predicting future observations, but dense latent-state propagation and sample-intensive trajectory optimization incur high inference latency and peak memory usage, limiting real-time deployment on resource-constrained platforms. Existing sparse world-model acceleration met...

Chun-Zheng Li, Ze-Sheng Jia, Hong-Da Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TAO-DA: Towards Autonomous Operation--A Dual-Arm Vision-Language-Action Model for Coordinated Manipulation

A symmetric Dual-Arm Expert (DAE) architecture built upon a shared Vision-Language Model (VLM) backbone with decoupled, arm-specific expert towers is proposed, providing preliminary evidence of emergent skill generalization from single- to dual-arm tasks (as well as the reverse), together with cross-arm motion-domain s...

Yong-Shen Zhao, Han Gao, Bao-Ping Cheng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Dynamic Manipulation with World-Action Models via Counterfactual Planning

DPP enables real-time dynamic manipulation on a single consumer GPU without additional training on dynamic data and constructs a counterfactual observation that places a predicted target position in a familiar robot context, allowing the model to invoke an existing manipulation skill rather than generate a recovery beh...

Sunwoo Park, Won-Sang Lee, Seonghyun Jin et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Tasks: A Vision for Reproducing an Animal-like Behavioral Substrate Using Modern Robot Learning Techniques

This work argues that their continual coordination under competing demands constitutes an important and underexplored target for modern robot learning and proposes the ethological behavioral substrate as a conceptual lens for studying this form of competence in artificial agents.

Samiyuru Menik, H. Jayalath · 0 citations
#artificial intelligence Preprint Sep 2026

VPTwin: Real-Sim-Real Video Prediction for Robotic Manipulation Planning

VPTwin is proposed, a Real-Sim-Real video prediction framework that anchors real-world future prediction using real-synchronized simulation twins and establishes a predictive planning loop using VPTwin to visually verify VLM-proposed actions and guide reliable real-world execution.

Zheng-Hao Xiao, Min-Ting Pan, Nan-Tian He et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.