Rearranging objects in cluttered tabletop environments remains a long-standing challenge in robotics. Classical planners often generate inefficient, high-cost plans by moving objects individually and using fixed buffers, temporary spaces such as unoccupied tabletop regions or static stacks, to resolve conflicts. When o...
Arman Barghi, Hamed Hosseini, Seraj Ghasemi et al.· 0 citations
This paper introduces POEF (POlicy EFfective Jailbreak), an automated red-teaming framework that takes into account the robot-specific constraints during both the optimization and evaluation processes and proposes two defense strategies that mitigate the behavior jailbreak risks.
Xuancun Lu, Zhen Huang, Xin-Feng Li et al.· 17 citations· ⚡5
A concept-centric framework for building agents that can learn continually and reason flexibly across multiple domains and offers several advantages, including data efficiency, compositional generalization, continual learning, and zero-shot transfer.
Jia-Yuan Mao, Joshua B. Tenenbaum, Jia-Jun Wu· Communications of the ACM· 14 citations· ⚡1
X-Reset is proposed, a framework that resolves exploration with human hand-object demonstrations of RL from scratch and scales with the number of training objects, generalizes to unseen objects, can learn from imperfect hand-pose estimates, and transfers behaviors zero-shot from sim-to-real.
Prithwish Dan, Chen-Yang Ma, Wei Zhan· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Failure for Rising is proposed, a failure-driven real-to-sim-to-real closed-loop learning framework that converts real-world failures into targeted policy improvement and reconstructs each failure as an interactive, object-centric table-top environment that preserves the task-relevant spatial and physical conditions.
Zhuo-Yuan Yu, Jia-Cheng Wang, Tian-Le Liu et al.· 0 citations
CATok is proposed, a causal action tokenizer that reframes tokenization as a causally structured generative process, establishing a high-performance, scalable foundation for purely autoregressive VLA systems.
Chen-Yu Zhang, Yu-Hang Cao, Da-Ru Du et al.· 0 citations
P2P-T is introduced, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations and achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently rema...
Bang-Jun Wang, Long-Yan Wu, Yu-Kun Wei et al.· 0 citations
Spatial Grafting is proposed, a versatile, lightweight spatial module that binds frozen reconstruction features to metric, robot-relative geometry and injects them into the flow-matching action expert through cross-attention without modifying the host's perceptual pathway, so the host retains the full benefit of its pr...
Ding-Sheng Liu, Yang-Zheng Wu, Mahboubeh Asadi et al.· 0 citations
Comparisons within ReCAT show that the observation encoder and every-block memory conditioning are needed for this performance, and that update rules developed for efficient sequence modeling behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule up...
Pankhuri Vanjani, M. Hatab, Can Mizrakli et al.· 0 citations
EMPIRIC, an agent that learns a residual world model: a physics engine extended with code for the missing mechanisms, that learns interpretable, reusable models, and solves more tasks with fewer environment interactions than all three baselines.
Yi-Chao Liang, Amber Li, Dat Nguyen et al.· 0 citations
MINT is proposed, which first trains VLA policies to remain functional under missing visual inputs, and selectively supplements missing observations using optical-flow extrapolation or an action-conditioned world model, and withdraws predicted views when they become unreliable.
Ming-Le Jiang, Rui Xu, Yun-Ke Wang et al.· 0 citations
RoboFL, which instantiates MoSAIC (Mixture of Slotted Adapters) for federated world-action learning, is presented, as it outperforms centralized PEFT InternVLA-A1 by 12.23% on the Franka arm, while reducing per-round client communication by up to 86.81% relative to MoE-based federated VLA baselines.
Rong-Yu Zhang, Rui-Zhi Fan, Yunfan Lou et al.· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.