HCI and HRI studies often require short, repeatable arousal manipulations that can run while participants continue interacting with a device or robot. These experiments are often challenged by the need to induce arousal in settings that still resemble real interaction. Participants must continue using a device, touchin...
Discrete action tokenization is central to autoregressive vision-language-action (VLA) models, yet action representations are often evaluated primarily through reconstruction fidelity. We ask which representation properties actually matter for closed-loop control by comparing fixed analytical, data-driven linear, and n...
Yu-Xin Yang, Gao-Han He, Chang-Xue Guan et al.· 0 citations
Robots operating safely in cluttered everyday environments often need to infer scene geometry from partial observations. Methods that detect objects in 2D and reconstruct them independently struggle in such scenes: a missed object is never reconstructed, a merged detection can fuse two objects, and separately reconstru...
Dongwon Son, Junhyek Han, Yoon-Je Cho et al.· 0 citations
Geometry-Change VLA (GC-VLA), which learns to predict multiview future-current geometry-change tokens from current observations and applies Geometry-Conditioned Residual Flow (GCRF), using a binary intervention router and a single bounded residual velocity policy learned from closed-loop feedback.
Jinu Pahk, Jesoon Kang, T. Park et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
We focus on human-robot collaborative transport, a challenging task of broad relevance spanning logistics, manufacturing, and the home, in which a user and a robot work together to relocate a large or heavy object. To act as an effective partner, the robot should reduce the user's effort by contributing to efficient re...
Neural inertial odometry increasingly uses networks as learned measurements inside filtering pipelines. Such measurements should transform consistently under arbitrary IMU mounting conventions: their mean must transform as a vector, and their covariance must transform congruently as a second-order tensor. We present GI...
Chan-Ki Kim, Ming-Han Zhu, Tzu-Yuan Lin et al.· 0 citations
What is the least recurrent memory needed to reproduce a specified expert under partial observability? The instantaneous requirement is the conditional entropy of the expert's behavioral quotient, but recurrence must also preserve distinctions that future observations will not restore before use. We characterize this m...
AgenticDiffusion is proposed, an agentic multi-view UAV navigation framework that semantically coordinates first-person-view and top-view observations for mission-level navigation that is robust to lexical variation in target descriptions.
Faryal Batool, Muhammad Ahsan Mustafa, Fawad Mehboob et al.· 0 citations
A closed-loop framework for autonomous magnetic microrobot navigation that separates long-range geometric planning from short-range reactive control is presented, and a modular framework for autonomous magnetic microrobot navigation in complex biological environments is established.
Yan-Da Yang, Max Sokolich, F. Kırmızıtaş et al.· 3 citations
Results show that structured agentic debugging can address a key cyber-physical integration bottleneck in real-world robot deployment, and are shown to be more complete and efficient recovery than human operators using Claude Code.
Minkyu Ham, Dongho Kim, Chan Lee et al.· arXiv.org· 0 citations
Self-driving laboratories (SDLs) are attracting increasing attention as a means of accelerating scientific discovery; however, developing SDL software remains technically demanding. To improve accessibility, orchestration software frameworks have been proposed to coordinate SDL components, but many existing frameworks...
Successful mobile manipulation requires coordinated base and arm motion while maintaining accurate spatial positioning. However, demonstration-trained policies can struggle to realise the intended base motion reliably, leading to spatial misalignment and subsequent manipulation failures. We present MAVP (Map-Aware Visu...
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.