TRWORLDBENCH is introduced, a benchmark for evaluating embodied world models through synchronized head, left-wrist, and right-wrist videos and uses 19 metrics to assess tri-view consistency, task alignment, physical and 3D coherence, motion quality, temporal consistency, and visual quality.
Xuan-Yi Liu, Hao-Feng Wang, Rui-Qi Li et al.· 0 citations
It is demonstrated that an attacker can embed a stealthy backdoor into an LLM-based robot controller by manipulating its instructions, triggered not by an external cue, but by a specific, rare sequence of the robot's own past actions.
Doniyorkhon Obidov, Shivayogi Akki, Cheng-Qiu Tan et al.· 2 citations
Large language models and vision-language models are increasingly used as high-level planners in robotic systems, using task goals and sensor summaries to select navigation or manipulation actions. This creates a new backdoor surface: a compromised planner can behave normally in most runs, yet change its target selecti...
Monocular drone navigation requires reaching a goal in an unseen environment from a single forward-facing camera, which offers few cues for depth and scale. World models address this by modelling how observations evolve under actions, but they are built to be executed: the prediction is produced at deployment and fed b...
Yu-Hang Zhang, Rangya Zhang, Yu-Jing Shang et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
DSR is introduced, a causal post-selection operator that repairs destination support without retraining the host predictor or increasing the maintained set size, and shows that finite-set support allocation is a useful prediction-side control point when a fixed hypothesis set serves as the interface to downstream syste...
Feng-Rui Liu, Jia-Jun Peng, Duo Peng et al.· 0 citations
Lower cost open source robots and reinforcement learning (RL) simulation tools create new opportunities for precollege students to engage with contemporary robotics. However, translating a complete research workflow, spanning mechanical assembly, electrical setup, simulation, policy learning, system identification, and...
IndustrialVLA-Bench is presented, an evidence-aware evaluation of six released VLA and WAM systems under a unified reporting schema that evaluates clean capability on LIBERO, non-language robustness on LIBERO-Plus, instruction sensitivity on LIBERO-Para, and observed execution cost.
Yi-Qi Wang, Zhi-Feng Rao, Jia-Qi Zhang et al.· 0 citations
These results establish recipe-dependent interactions and identify concrete precision assignments, rather than universal layer-sensitivity rules, rather than universal layer-sensitivity rules for vision-language-action models.
Jiu-Yi Xu, Qing Jin, Mei-Da Chen et al.· 0 citations
Ubiquitous robotic systems often lack traditional visual interfaces, necessitating resilient natural language interaction for maintenance and repair tasks. This paper presents a goal oriented agentic AI architecture designed to enable non-expert users to perform technical repairs through situated dialogue. The framewor...
Agriculture 4.0 robotic systems improve field efficiency yet remain too capital-intensive for the fragmented smallholdings that dominate global agriculture. Meanwhile, a growing number of retired low-speed electric-vehicle (LSEV) powertrains retain functional electromechanical value but are destructively recycled. This...
Weijie Shi, Zicheng Xu, Zhenbang Cheng et al.· 0 citations
World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to calibration, perception, and contact-dynamics errors in real-world precision tasks, often failing in the final few millimeters of alignment or insertion. We propose HALO-WA, a hybrid-...
Angen Ye, Weijie Ke, Xiaofeng Wang et al.· 0 citations
Neural world models coupled with model predictive control (MPC) replan at every environment step to bound accumulated prediction error, but this incurs substantial computational overhead. Reusing a cached plan reduces this overhead, yet its effectiveness depends on how prediction mismatch propagates through the local d...
Yutian Cheng, Xiaojian Ma, Xianhao Wang et al.· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.