Aerial manipulators extend robotic manipulation into 3D workspaces that are difficult for ground-based robots to access, creating new opportunities for general-purpose manipulation. However, extending Vision-Language-Action (VLA) models to aerial robots introduces distinct challenges due to the tight coupling between m...
Rui Huang, Yan-Lin Mu, Li-Dong Li et al.· 0 citations
Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or determi...
Junhyun Ha, Ju-Ho Lee, Byoungwoo Park· 0 citations
Spotter is proposed, which reverses the roles: the embodied model leads and executes continuously, while the VLM runs in parallel, monitors through a lightweight local screener, intervenes only when an error is detected, reflects on and corrects it, and returns control.
Long Li, Qi-Chao Zhao, Yue Yang et al.· 0 citations
This work studies what determines whether predictive supervision improves the visual representation used by a VLA policy, and finds that different prediction interfaces produce markedly different forecasts and visual representations, including in the spatial, dynamics, and action information that transfers beyond famil...
Hanseul Kim, J. Yeom, Youngjoon Jeong et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Visual-memory systems commonly retain or compress past observations. Robot control additionally requires interaction-derived state that no individual frame may explicitly represent, such as persistent identity relations, accumulated progress, or ordered procedures. We introduce Simple Agentic Robot Memory (SimpleARM),...
Yu-You Zhang, Yun-Bei Zhang, Miao Li et al.· 0 citations
We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (S...
Rui-Xiao Xu, W. Kenny, Zhi-Qian Liu et al.· 0 citations
Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introdu...
Yun-Bei Zhang, Zi-Jian Jin, Yuan-Zhe Liu et al.· 0 citations
STAIRCASE POLICY is introduced, a streaming inference and training framework that turns a flow-matching VLA into a JEPA-style WAM and partitions a large action chunk into sub-chunks at staggered denoising stages, enabling long-horizon execution without repeated full policy inference.
Guoheng Sun, Chen Chen, Jin Wang et al.· 0 citations
FineART is presented, a densely annotated bimanual manipulation dataset comprising 40,543 episodes and 533,913 subtasks across 151 tasks and FineART-VLA is introduced, a vision-language-action policy that predicts its own next subtask to guide its actions.
Jade Choghari, Pepijn Kooijmans, Mansi Agarwal et al.· 0 citations
This work introduces Aligned Transport of Latent Structure (ATLAS), a training objective that explicitly preserves relational geometry while calibrating the global latent distribution and uses Wasserstein embedding matching to calibrate its marginal through one-dimensional Wasserstein-2 transport.
A large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task, and a unified evaluation framework with a unified evaluation framework.
Tong-Tong Feng, Xin Wang, Hao-Ran Hou et al.· 0 citations
Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowled...
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.