Skip to content

Category

robotics

1,156 papers

#artificial intelligence Preprint Open access Sep 2026

Dynamic Buffers: Cost-Efficient Planning for Tabletop Rearrangement with Stacking

Rearranging objects in cluttered tabletop environments remains a long-standing challenge in robotics. Classical planners often generate inefficient, high-cost plans by moving objects individually and using fixed buffers, temporary spaces such as unoccupied tabletop regions or static stacks, to resolve conflicts. When o...

Arman Barghi, Hamed Hosseini, Seraj Ghasemi et al. · 0 citations
#artificial intelligence Preprint Dec 2024

Easier Said Than Done: Unpacking Intent-Behavior Gap in Jailbreaking LLM-Based Robots

This paper introduces POEF (POlicy EFfective Jailbreak), an automated red-teaming framework that takes into account the robot-specific constraints during both the optimization and evaluation processes and proposes two defense strategies that mitigate the behavior jailbreak risks.

Xuancun Lu, Zhen Huang, Xin-Feng Li et al. · 17 citations · ⚡5
#artificial intelligence Open access May 2025

Building Intelligent Agents with Neuro-Symbolic Concepts

A concept-centric framework for building agents that can learn continually and reason flexibly across multiple domains and offers several advantages, including data efficiency, compositional generalization, continual learning, and zero-shot transfer.

Jia-Yuan Mao, Joshua B. Tenenbaum, Jia-Jun Wu · 14 citations · ⚡1
#artificial intelligence Preprint Sep 2026

X-Reset: Scaling Object-Centric Reinforcement Learning via Cross-Embodiment Resets

X-Reset is proposed, a framework that resolves exploration with human hand-object demonstrations of RL from scratch and scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.

Prithwish Dan, Chen-Yang Ma, Wei Zhan · 0 citations
#artificial intelligence Preprint Sep 2026

F4R: Failure-Driven Recognition, Reconstruction, Refinement, and Redeployment for Continual Robot Self-Improvement

Failure for Rising is proposed, a failure-driven real-to-sim-to-real closed-loop learning framework that converts real-world failures into targeted policy improvement and reconstructs each failure as an interactive, object-centric table-top environment that preserves the task-relevant spatial and physical conditions.

Zhuo-Yuan Yu, Jia-Cheng Wang, Tian-Le Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations

P2P-T is introduced, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations and achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently rema...

Bang-Jun Wang, Long-Yan Wu, Yu-Kun Wei et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Spatial Grafting: Grounding 3D Features for Flow-Matching Robot Policies

Spatial Grafting is proposed, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry and injects them into the flow-matching action expert through cross-attention without modifying the host's perceptual pathway, so the host retains the full benefit of its pr...

Ding-Sheng Liu, Yang-Zheng Wu, Mahboubeh Asadi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ReCAT: Remember, Count, and Time: Structured Recurrent Memory for Robot Manipulation

Comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance, and that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule up...

Pankhuri Vanjani, M. Hatab, Can Mizrakli et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Learning to Act under Visual Interruptions with Vision-Language-Action Models

MINT is proposed, which first trains VLA policies to remain functional under missing visual inputs, and selectively supplements missing observations using optical-flow extrapolation or an action-conditioned world model, and withdraws predicted views when they become unreliable.

Ming-Le Jiang, Rui Xu, Yun-Ke Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

RoboFL: Federated Expert Assembly for World Action Models

RoboFL, which instantiates MoSAIC (Mixture of Slotted Adapters) for federated world-action learning, is presented, as it outperforms centralized PEFT InternVLA-A1 by 12.23% on the Franka arm, while reducing per-round client communication by up to 86.81% relative to MoE-based federated VLA baselines.

Rong-Yu Zhang, Rui-Zhi Fan, Yunfan Lou et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.