Progress-Heuristicized Inverse Reinforcement Learning (PHIRL), a data-efficient framework that learns robust reward functions by jointly leveraging demonstrations and feedback, uses progress, a feedback modality that describes cumulative task completion.
Hang Yu, James Staley, Cheng-Xi Tsou et al.· 0 citations
The results do not establish generalization to other lots or real vehicles, or a formal safety guarantee, and the results do not establish generalization to other lots or real vehicles, or a formal safety guarantee.
Ke-Jia Gao, Li-Guo Zhou, Lei Yu et al.· 0 citations
A practical way to build general-purpose manipulation experiments around GPT-6-Astra is suggested: supply machine-readable body descriptions and synchronized demonstrations, and turn useful agent-generated feedback routines into reusable skills, while the agent adapts actions from current images.
Si-Da He, Ling-Xi Xie, Yun-Ning Cao et al.· 0 citations
Can a robot improve itself the way coding agents now improve software? We built an agentic system to find out. It watches a robot fail, works out which capability is missing, writes new skills or finds and installs external models, tests every change in simulation, and repeats, with no human writing robot code. We ran...
Jiaming Wang (National University of Singapore)· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
JRDB-AVR is introduced, a benchmark derived from existing real-world JRDB robotics data through a structured question-generation engine that turns this gap between answer accuracy and evidence accuracy in current baselines into an explicit evaluation.
Zhixi Cai, Fu-Cai Ke, Sukai Huang et al.· 0 citations
Vision-language-action (VLA) models adapt pretrained vision-language models (VLMs) for closed-loop robot control, transferring their perceptual and semantic capabilities to action prediction. Despite strong in-distribution performance, however, VLAs often degrade under deployment shifts. Adversarial training (AT) offer...
Signal Temporal Logic (STL) enables rigorous verification and control of cyber-physical systems, but writing correct specifications requires expertise that most requirement holders lack. Large language models can translate natural-language (NL) requirements into STL, yet stronger translators alone approach an accuracy...
This work introduces PlanGuard, the first pre-execution detector that evaluates the physical safety of a complete multi-step plan in its current environment, and proposes Strong-Teacher Adaptive Compensation for On-Policy Distillation (STAC-OPD), which provides compact models with adaptive strong-teacher supervision al...
Jun-Chi Chen, Chang-Tao Miao, Yu Xiang et al.· 0 citations
Understanding the safety risks of vision-language-action (VLA) models is essential for their deployment in the physical world. Existing safety research has mainly considered persistent perturbations that are applied continuously to observations throughout an episode. However, momentary observation corruption, in which...
Model-predictive control with Joint-Embedding Predictive Architectures (JEPAs) provides a strong zero-shot goal-reaching planner, but it is only effective over short planning horizons. Hierarchical extensions attempt to bridge this gap by learning a macro planner to predict intermediate latent sub-goals to guide the mi...
Royson Lee, Fady Rezk, Titouan Parcollet et al.· 0 citations
While Vision-Language-Action (VLA) models perform strongly on manipulation tasks, their responses to invalid task premises remain underexplored. Existing evaluations of premise conflicts often focus on terminal task outcomes, yet task failure alone cannot distinguish behavioral disengagement from continued pursuit foll...
Teleoperated ultrasound can improve diagnostic medical imaging access for remote communities. Having accurate force feedback is important for enabling sonographers to apply the appropriate probe contact force to optimize ultrasound image quality. However, large time delays in communication make direct force feedback im...
Ryan S. Yeung, David G. Black, Septimiu E. Salcudean· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.