This work proposes Action Upcycling, a training-free algorithm that reuses actions the policy would otherwise discard, without accessing model internals or drawing extra samples, and finds that discarded actions stay close to their replanned versions as long as the action velocity remains smooth.
Taesung Kwon, Jangho Park, Sunwoo Park et al.· 0 citations
Many physical tasks in human environments require collaboration, from assisting a partner to jointly manipulating an object. Yet, existing humanoid benchmarks largely focus on single-humanoid skills and lack evaluation of multi-humanoid collaboration under egocentric visual observations. We introduce CoHuB (Collaborati...
Hyunjin Park, Jebeom Chae, Minwoo Park et al.· 0 citations
Reward shaping is fundamental to modern robotic control with deep reinforcement learning (RL), yet practitioners still rely heavily on heuristic principles borrowed from classical optimal control and trajectory optimization. Existing methods rarely distinguish reward terms that are intrinsic to the control objective fr...
World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and interme...
Zhi-Nan Liu, Hao-Zhi Han, Ru-Ge Zhang et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work introduces LRC-JEPA, a lightweight end-to-end world model that routes information into a compact predictive latent and learned-query residual-context embeddings and shows that the resulting representation is sufficient, minimal, nuisance-invariant, and disentangled.
Lu-Zhe Huang, Lei Chu, Jing-Yi Liang et al.· 0 citations
A framework termed Predictive Semantic Safety (PSS), which connects visual physical reasoning to backup-based safety filtering and derives input-affine constraints for minimally modifying the nominal input while preserving backup feasibility under the robot dynamics and input limits.
Taekyung Kim, S. Fradi, Yan-Ning Dai et al.· 0 citations
World models offer a powerful substrate for safety reasoning in high-dimensional robotic systems, but they are also fallible: their predictions can be biased, miscalibrated, or confidently wrong. This creates a central challenge for latent-space safety filters, which often learn Hamilton-Jacobi safety value functions o...
This work proposes a Bayesian framework that treats clarification as an active learning problem over grounded Signal Temporal Logic task specifications and uses LLMs to initialize candidate formal specifications and translate informative contrasts into natural-language clarification questions, while Bayesian optimizati...
Hu-Ao Li, Carson Sobolewski, A. Saravanos et al.· 0 citations
SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and unt...
T. Bui, Jongsul Moon, Youngouk Kim et al.· 0 citations
Pretrained world action models (WAMs) provide generalist capabilities across diverse robotic manipulation tasks, yet improving target-task performance to an expert level without degrading pretrained skills remains challenging. We explore on-policy distillation (OPD) for WAMs and introduce WAM-OPD. WAM-OPD inherits the...
P. Liu, Xiao-Han Lei, Shi-Qi Zhang et al.· 0 citations
Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated b...
Wen-Qiao Li, Qian-You Zhao, Jia-Wen Hao et al.· 0 citations
Autonomous driving requires \textit{world models} that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM for end-to-end auto...
Hao-Ran Zhu, Wan-Cong Zhang, Yann LeCun et al.· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.