Skip to content

Category

robotics

1,156 papers

#artificial intelligence Preprint Sep 2026

Make Code as Policy Great Again: Frontier Agents Write, Call, and Evolve Robot Tools

Frontier models can control robots, but reasoning through every reach, grasp, and retreat makes manipulation slow and token-intensive. We revisit code as policy with a different division of labor: models build executable tools, code handles multi-phase motions, and models decide what to do next. We introduce URAI (Univ...

S. Ge, Alex Zhou, Jian-Shu Zeng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Sparse-WAM: Accelerating World Action Models via Action-Guided Sparse Imagination

World-action models (WAMs) leverage pretrained video models to improve generalization in robot control by jointly predicting future visual states and actions. This capability comes at a substantial inference cost, as dense future-frame tokens are repeatedly processed during denoising. Prior methods address this by toke...

Xin-Ling Xie, Hao-Dong Wang, Jia-Zhi Mi et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SimEX: Simulation-Integrated Robotics AutoResearch

Coding agents powered by large language models (LLMs) have shown remarkable abilities to autonomously reason about and achieve goals in the digital world. However, bringing this success to the physical world remains challenging. On the one hand, direct generation methods (e.g., Code as Policies) often suffer from the L...

Jiaheng Hu, Roberto Martin-Martin, Peter Stone et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DrivingBench: Can Vision-Language Models Drive a Toyota Corolla?

Frontier models excel at many digital benchmarks, yet their ability to drive a real car, an everyday human skill, remains largely untested. We present DrivingBench, to our knowledge the first benchmark where general-purpose vision-language models must drive a real car. Through three tools, the models see camera frames...

Aditya Ramabadran, Simon Mahns, Tobias Gessler · 0 citations
#artificial intelligence Preprint Sep 2026

Efficient Multi-Modal Planning with Reward-Guided Preference Optimization for Autonomous Driving

Safe and efficient trajectory planning is essential in autonomous driving. However, existing end-to-end approaches often fall short in both computational efficiency and safety guarantees. Methods based on imitation learning suffer from causal confusion, while rule-based scoring approaches often incur heavy computationa...

Cheng-Lin Chen, Lu-Jia Wang, Xin-Hu Zheng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Online Evolution Strategy for Flow-Matching VLA Policies via Self-Supervised Trajectory Distribution Optimization

Vision-Language-Action (VLA) models based on generative frameworks, such as Flow Matching, have recently achieved impressive performance in robotic manipulation. Unlike deterministic policies, Flow Matching enables VLA models to learn conditional action trajectory distributions, where latent noise vectors induce differ...

Gongxin Yao, Yongsheng Zhao, Jiayin Deng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

TERRA: Terrain-Aware Reconstruction, Retargeting and Control for Musculoskeletal Locomotion

Recent advances in musculoskeletal modeling and reinforcement learning have enabled muscle-actuated agents to reproduce increasingly complex human motions. Yet these capabilities remain largely confined to flat ground, in part because motion datasets rarely include aligned terrain geometry and because retargeting terra...

Merkourios Simos, Chengkun Li, Bianca Ziliotto et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

GestAdapt: Workspace-Conditioned Co-Speech Gesture Generation for Humanoid Robots

Co-speech gestures for robots must adapt not only to speech and embodiment, but also to the workspace available for performing the motion. Since the same speech can be accompanied by different gestures, a robot can respond to workspace constraints, e.g., gestures for speech next to a wall. In these scenarios, the robot...

Bosong Ding, Xianglin Zhang, Miao Xin et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

SynIL: Leveraging Synergy for Offline Imitation Learning from Imperfect Demonstration Datasets

Imitation learning enables robots to acquire complex skills directly from massive demonstration datasets, but its performance degrades severely when datasets are contaminated with suboptimal or noisy demonstrations. While prior quality-assessment methods attempt to filter or reweight data, they typically rely on manual...

Yuto Tanaka, Kyo Kutsuzawa, Martina Doku et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Belief-Aware Multi-Agent Path Finding under Map Uncertainty

Multi-Agent Path Finding (MAPF) aims to find collision-free paths for multiple agents in a shared environment. Classical MAPF assumes that all static obstacles are known in advance, but real-world environments can change unexpectedly due to fallen objects, spills, or other local disturbances. When such changes are spat...

Viraj Parimi, Shao-Hung Chan, Han Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Divide and Collapse: MAPF-Collapse via Exact Decomposition into Independent Sub-Instances

In this work we study the problem of MAPFC, a post-optimization step for Multi-Agent Path Finding (MAPF) plans where we are given a feasible plan produced by a modern MAPF solver and are tasked with removing avoidable moves while preserving feasibility. This NP-hard problem naturally arises when using learning-based st...

Oren Salzman · 0 citations
#machine learning Preprint Open access Sep 2026

TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback

Contact-rich manipulation requires adapting to contact states that can evolve substantially within an action horizon. However, chunk-based vision-language-action models predict complete action chunks from observations collected before execution, leaving tactile conditioning stale during execution. Existing tactile-reac...

Jianbo Zhou, Boyuan Zhao, Yuzheng Zhang et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.