The Agentic Reward System (ARS) is introduced, an inference framework for progress reward modeling with general-purpose vision-language models (VLMs), without additional reward-model training, and its results suggest that structured inference and verification can improve the usefulness of general-purpose VLMs for robot...
Sheng-Miao Hu, Wei-Yi Lu, Ling-Bing Zeng et al.· 0 citations
World models for manipulation are typically trained from video, yet the events that determine how manipulation unfolds, such as making and releasing contact, are difficult to observe visually and are often easier to sense through touch. We present the Dexterous Tactile World Model (DTWM), a video world model for future...
Zi-Yao Zeng, Xiatao Sun, Hao Wang et al.· 0 citations
General-purpose vision-language models offer a promising way to zero-shot robot control: \gptastra{} excels at open-ended and language- or image-conditioned manipulation but remains substantially weaker on high-precision and long-horizon tasks. We introduce \emph{RoboICL}, an in-context robot-control framework that nar...
Fang-Cheng Liu, Ye-Qing Shen, An-Da Cheng et al.· 0 citations
This paper proposes a 3D point tracker accurate in those absolute terms and operating within a single commodity GPU, pose-free, monocular budget, exceeding strong feed-forward trackers and a companion analysis explains why several published trackers lose most of their accuracy under this budget.
Masahiro Ogawa, Qi An, Atsushi Yamashita· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Simulation-trained manipulation policies can exploit privileged state information to learn effective contact-rich behaviours, but deployment requires acting from partial observations such as noisy camera images. A common solution is teacher-student distillation, in which a visuomotor policy is trained to reproduce the...
Denis Shcherba, Adrian Abel, Eckart Cobo-Briesewitz et al.· 0 citations
This work decomposes multi-step rollout error into the errors introduced at individual steps and their propagation through subsequent transitions, showing that state-affine dynamics are precisely the differentiable transitions with state-independent Jacobians, eliminating the nonlinear propagation residual and making t...
Bo-Yuan Zhang, Ying-Jun Du, Xian-Tong Zhen et al.· 0 citations
This work evolves freeform robots to pick up, hold, rotate, and use diverse objects, and uses contrastive learning to create a highly searchable genetic embedding of design space, an autoregressive developmental model to decode designs, evolutionary strategies to find good designs, and reinforcement learning to train e...
Zihan Guo, Shu-Zhe Zhang, Mu-Han Li et al.· 0 citations
A graph-attention network that reads the current and proposed states of moving agents together with locally relevant stationary obstacles and returns one collision-risk score per moving agent, which any controller can use to accept, repair, replan, or postpone a proposed step.
Alan Debbas, Edwin Meriaux, Gregory Dudek· 0 citations
Copper-Policy is introduced, which learns a compact World representation with the policy rather than relying on a predefined target space and predicts future observation embeddings conditioned on task intention without reconstructing pixels through temporal joint-embedding prediction.
Ze Feng, Yi-Xu Feng, Ling-Yu Xiao et al.· 0 citations
Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations, making deployment in diverse environments a challenge. We present a controlled empirical study of which design choic...
Mino Nakura, Sriram Krishna, Yufei Wang et al.· 0 citations
This work proposes Fast-TD-MPC, a lightweight framework that adaptively routes between fast policy execution and test-time planning, reserving costly deliberation for states where it is most needed, and delivers competitive task performance across 103 continuous control tasks while achieving up to ~4x faster inference.
Federated learning offers a natural way for multiple robots to jointly improve manipulation policies without requiring centralized access to training demonstrations. However, non-IID task and environment distributions can induce representation drift and mutually incompatible robot-policy updates, making naive parameter...
Biprodip Pal, Kaushik Roy, Yan-Ming Zhu et al.· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.