Skip to content

Category

robotics

1,156 papers

#artificial intelligence Preprint Oct 2026

Managing Context and Communication in Distributed Agentic UAV Swarms

Unmanned aerial vehicle (UAV) swarms increasingly rely on language-model agents to provide adaptive mission-level reasoning in uncertain environments. Fully distributed control, in which each UAV hosts an independent Small Language Model (SLM), removes reliance on a centralized coordinator but introduces an information...

Andrea Iannoli, Ivan D. Zyrianoff, A. Trotta et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Completion Aware Guidance for World Action Models

World Action Models (WAMs) predict visual futures and robot actions, yet they remain susceptible to task-incomplete imagination, where plausible, action-consistent predictions omit the transition needed for task completion. In this paper, we show that this failure is not inherent to the world model backbone, but emerge...

Seungyeon Kim, Junhoo Lee, Baekseung Kim et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots

Quadrupedal locomotion requires balancing conflicting objectives such as command tracking, stability, and energy efficiency, yet conventional reinforcement learning (RL) hardcodes these priorities into a fixed scalar reward at training time. We present PROMO (Preference-Conditioned Multi-Objective Reinforcement Learnin...

Amr Mousa, Rifny Rachman, Neil Karavis et al. · 0 citations
#artificial intelligence Preprint Oct 2026

OrbitTAMP: Grounding Language Models for Task and Motion Planning in Spacecraft Rendezvous

A hierarchical framework for spacecraft task-and-motion planning (TAMP) that grounds LLM reasoning in a graph of reusable behaviors and domain-specific planning modules is presented, establishing a scalable and auditable foundation for language-driven agentic planning of spacecraft RPO.

Yuji Takubo, Daniele Gammelli, Marco Pavone et al. · 1 citation
#artificial intelligence Preprint Open access Oct 2026

Divide-and-Remember: Recursive Action-Relevant Memory for Long-Horizon VLA Policies

Vision-language-action (VLA) models struggle on history-dependent manipulation tasks, where the current observation alone does not determine the action, and the policy needs a memory of the history. Existing memory methods decide what to remember by design, for example, keeping frames with large pixel changes, and show...

Xuehui Yu, Eason Yu, Meiyi Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Screw Attention: Rigid-Body Algebra Inside a Transformer

Learned manipulation policies rediscover from data the spatial relations that rigid-body mechanics supplies in closed form. This costs data, and it leaves the policies fragile to geometric changes in the scene. We present Screw Attention, a transformer layer in which the relation between two bodies is a spatial transfo...

Aly Magassouba · 0 citations
#artificial intelligence Preprint Oct 2026

TOAST: Stochastic Robot Action Tokenization for Autoregressive Vision-Language-Action Models

Autoregressive Vision-Language-Action models often represent continuous robot actions as discrete token sequences, enabling action prediction with standard next-token objectives. FAST has substantially improved this representation by compactly encoding action containing diverse temporal frequencies into relatively few...

Keisuke Shirai, Tomohiro Motoda, Hanbit Oh et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models

In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a promising framework for this goal, yet its direct application leads to performance collapse. We discover that this stems f...

Jiawei Fan, Sifeng Wang, Yuqing Hou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Learning Transferable Skills using Goal-Conditioned Bisimulation

Unsupervised skill discovery has emerged as a promising approach for leveraging reward-free datasets to pretrain general-purpose policies. However, current skill discovery methods either require access to expert data or exhibit limited generalization, failing to transfer effectively to previously unseen layouts. A key...

Mohammad Amin Abbasfar, Farbod Azimmohseni, M. Rohban · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MIKASA-Robo-VLA: Benchmarking Memory in VLA Models for Long-Horizon Manipulation

Vision-language-action policies often see only one or a few recent frames, which makes it difficult to evaluate how they use information that disappears during a task. We introduce MIKASA-Robo-VLA, a benchmark of 90 language-conditioned manipulation tasks. All but 10 hide the cue an action depends on. Those 10 are reac...

Egor Cherepanov, Nikita Kachaev, Aleksandr I. Panov et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies

Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface for runtime safety through reasoning monitoring and correction. In this work, we define and operationalize two evaluation axes for assessing when this interface can improve...

Sathwik Karnik, Joseph JR. Lee, Aryaman Gupta et al. · 0 citations
#artificial intelligence Preprint Sep 2026

DeepJEPA: Scaling World Models from Within

World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. We show that making every transition uniformly deeper wastes computation and can degrade planning because useful refinement is concentrate...

Zi-Jian Jin, Yun-Bei Zhang, Yuan-Zhe Liu et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.