Skip to content

Category

robotics

1,156 papers

#machine learning Preprint Open access Oct 2026

Flash-WAM: Modality-Aware Distillation for World Action Models

World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control. Step distillation has emerged as the natural remedy, but off-the-shelf methods b...

Arman Akbari, Ci Zhang, Arash Akbari et al. · 0 citations
#machine learning Preprint Open access Oct 2026

DyQ-VLA: Temporal-Dynamic-Aware Quantization for Embodied Vision-Language-Action Models

Vision-language-action (VLA) models achieve favorable task performance, yet runtime errors in closed-loop execution evolve with alternating updates of actions and observations. Existing VLA quantization methods mainly trade off inference speed and model performance while rarely investigating how quantization alters run...

Zihao Zheng, Hangyu Cao, Chibang Tao et al. · 0 citations
#machine learning Preprint Open access Oct 2026

How Vulnerable Is My Learned Policy? Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies

Imitation learning, also known as learning from demonstrations, is a popular approach to train AI models; however, the vulnerability of these models to adversarial attacks remains underexplored. We present the first systematic study of adversarial attacks, across a range of both classic and recently proposed imitation...

Akansha Kalra, Basavasagar Patil, Guanhong Tao et al. · 0 citations
#machine learning Preprint Open access Oct 2026

On Learning Optimal Corners in Orthogonal Partially Observable Cooperative Guard Art Galleries

The CADENCE algorithm solves the Partially Observable Cooperative Guard Art Gallery Problem (POCGAGP) with formal coverage and connectivity guarantees, but leaves unspecified which valid corner each agent should be deployed to, a choice that strongly affects efficiency. We introduce two learned corner-selection heurist...

Yassin Ben Mansour, Edwin Meriaux · 0 citations
#machine learning Preprint Oct 2026

RealtimeWAM: One-Step Asynchronous World Action Models

World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action...

Cheng-Tao Lv, Jin-Yang Du, Shu-Yi Feng et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Dual Variational Autoencoders for Efficient Sim-to-Real Transfer in Low-Cost Robotic Navigation

Vision-based autonomous navigation for low-cost robots remains a fundamental challenge, primarily due to the significant gap between simulated training environments and real-world operational conditions. Direct policy transfer from simulation is often ineffective, while training exclusively on real data is impractical....

\'Alvaro D\'iez (Department of Computer Science, Artificial Intelligence, University of Alicante) et al. · 0 citations
#machine learning Preprint Oct 2026

Encoded but Not in Control: Revealing the Grounding Gap in Vision-Language Robot Policies

Instruction following is central to language-conditioned robot policies: language should determine what to do when the same scene permits multiple valid actions. Yet successful execution alone cannot establish whether a policy follows the instruction or infers the task from the scene. We study this ambiguity through sc...

Shao-Han Jiang, Jia-Hang Cao, Qi-Duo He et al. · 0 citations
#machine learning Preprint Open access Oct 2026

AUTOPILOT An Advanced Perception, Localization and Path Planning Techniques for Autonomous Vehicles Using YOLOv7 and MiDaS

Self driving vehicles have emerged as a reliable technology that has the capability to transform transportation and mobility. The development of self driving cars requires significant advances in a number of areas, including perception, localization, decision making, and control. This research paper is based on the pro...

Harshkumar Devmurari, Gautham Kuckian, Prajjwal Vishwakarma · 0 citations
#machine learning Preprint Open access Oct 2026

P3: Persistent Particle Planning for Constrained Diffusion Control

Diffusion models provide expressive priors over trajectories, but adapting these priors to test-time constraints requires maintaining feasibility and consistency across successive control decisions. We introduce Persistent Particle Planning (P3), a sequential Monte Carlo framework for diffusion control that maintains a...

Hikmet Simsir, Mahyar Fardinfar, Ozgur S. Oguz · 0 citations
#machine learning Preprint Open access Oct 2026

How (and How Not) to Use Data Augmentation in VLA Post-Training

Vision-language-action (VLA) models currently demonstrate strong performance in a wide range of real-world robotics tasks. However, they often still lack the generalization ability to handle large visual out-of-distribution shifts. Post-training of VLAs with reinforcement learning (RL) has been shown to benefit robustn...

Bram Grooten, Joaquin Vanschoren · 0 citations
#machine learning Preprint Oct 2026

Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning

Learning from human demonstrations is a reliable way to teach robots new tasks, but the gains from each additional demonstration shrink as the policy improves. Continued improvement can instead come from supervised deployment, where an operator places the objects and intervenes when the policy fails. We ask how to maxi...

Lars Ankile, Perry Dong, R. Bhowmik et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Transporting Unsecured Stacked Payloads with a Quadrupedal Robot via Multi-Objective Reinforcement Learning

Transporting unsecured payloads with legged robots over uneven terrain requires balancing locomotion performance and payload stability, since aggressive motion can destabilize the payload even when the robot remains stable. We study quadrupedal transportation of unsecured stacked boxes on an edgeless torso-mounted boar...

Nobuo Namura, Masayuki Hiromoto, Kento Uemura et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.