Skip to content

Category

robotics

1,156 papers

#robotics Preprint Open access Sep 2026

Robot Programming with Augmented Reality: The Role of Spatial Ability

Programming a robot arm requires users to interpret coordinate frames, joint rotations, and trajectories that are not directly visible. Augmented reality (AR) can make these spatial relations visible, but its benefits may depend on users' spatial ability. We conducted a randomized between-subjects experiment ($N=71$) i...

Nicolas Leins, Muriel Fischer, Malte Teichmann et al. · 0 citations
#robotics Preprint Open access Sep 2026

ERUPT: An Open Toolkit for Interfacing with Robot Motion Planners in Extended Reality

We present the Extended Reality Universal Planning Toolkit (ERUPT), an extended reality (XR) system for interactive motion planning. This paper serves to introduce our open-source system to others who can use it as a base to develop immersive robot interaction applications. Our system allows users to create and dynamic...

Isaac Ngui, Courtney McBeth, Andr\'e Santos et al. · 0 citations
#robotics Preprint Open access Sep 2026

Ask Before It Tells: Benchmark-to-Robot Body-Cue Transfer for a Question-First Bedside Robot

Body-cue recognition can support assistive robots, but benchmark accuracy does not guarantee reliable behavior under a robot-camera viewpoint. We present Nuni, a bedside robot prototype that treats a detected distress cue as a reason to ask rather than a reason to alert. We compare two X3D-UGT RGB appearance classifier...

Dongsik Yoon · 0 citations
#robotics Preprint Sep 2026

Toward Human-in-the-Loop Robot Failure Recovery: Bridging Communication Gaps in Human-Robot Collaboration

Robots can recover from failures by asking bystanders for help, but effective human-in-the-loop recovery requires communication that accounts for differences in people's knowledge. Prior inverse-semantics work generates requests using a single listener model, leaving differences in listener knowledge untested. We intro...

Promise Ekpo, T. Vijay, Dhruv Mandalik et al. · 0 citations
#robotics Preprint Open access Sep 2026

Identity Continuity in Long-Term Embodied AI Relationships: From Agent-Specific Identity Representation to Identity-Continuity Appraisal

Long-term embodied AI will undergo learning, model updates, memory compression, hardware repair, and migration across embodiments. For users who have formed sustained relationships with such systems, these changes raise not only a problem of product consistency but also one of identity continuity: whether the changed s...

Zijian Ru · 0 citations
#computer vision Preprint Open access Sep 2026

3D-MoE: Towards Spatial Intelligence with Mixture-of-Experts for 3D Reasoning and Action Generation

Spatial intelligence, encompassing 3D perception and reasoning, is the essential next frontier of AI. Scaling current 3D vision-language models (VLMs) that rely on dense Transformers for spatial tasks incurs prohibitive computational costs. In this paper, we introduce 3D-MoE, a 3D VLM leveraging an efficient mixture-of...

Yueen Ma, Zenglin Xu, Irwin King · 0 citations
#artificial intelligence Preprint Open access Sep 2026

GraRe: Grasp Candidate Re-Ranking for Frozen 6-DoF Grasp Detectors

Existing 6-DoF grasp detectors typically rank grasp candidates by detector confidence. However, our analysis on GraspNet-1Billion shows that detector confidence is often poorly aligned with grasp quality, leaving successful grasp candidates at low ranks. Motivated by this observation, we study whether learned re-rankin...

Jibao Yuan, Yuhui Zhao, Yinzhen Lv et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)

I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 teams in the online (simulation) round and 2nd in the real-world final. It improves a vision-language-action (VLA) policy with a reinforcement-learning loop. The policy is its own value...

Ilia Larchenko · 0 citations
#artificial intelligence Preprint Open access Sep 2026

InSight: Self-Guided Skill Acquisition via Steerable VLAs

Vision-language-action (VLA) models excel at robot manipulation via imitation learning, but adapting them to new tasks often requires additional human demonstrations, which can be costly or infeasible. Meanwhile, vision-language models (VLMs) offer semantic task understanding but lack the physical grounding required fo...

Maggie Wang, Lars Osterberg, Stephen Tian et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models

Deploying vision--language--action (VLA) models on robots requires adapting model-specific inference pipelines to heterogeneous processors and limited onboard memory. We present vla.cpp, a unified C++ inference runtime for eleven VLA models, with no PyTorch dependency for model execution. The runtime shares model loadi...

Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen et al. · 0 citations
#machine learning Preprint Open access Sep 2026

PerchRL: Vision-Based Agile Perching on Inclined Platforms under Rapid and Irregular Motion

Autonomous vision-based perching of quadrotors on moving inclined platforms is critical for air-ground collaboration but remains challenging due to the limited field of view (FOV). In this paper, we propose PerchRL, a reinforcement learning (RL) framework for vision-based agile perching on inclined platforms under rapi...

Zihong Lu, Zongzhuo Liu, Huaxu Li et al. · 0 citations
#machine learning Preprint Open access Sep 2026

On Learning Spatial Structure from Pre-Beamforming Per-Antenna Range-Doppler Radar Measurements

Automotive radar perception pipelines commonly construct angle-domain representations via beamforming before applying learning-based models. This work instead investigates a representational question: can meaningful spatial structure be learned directly from pre-beamforming per-antenna range-Doppler (RD) measurements?...

George Sebastian, Philipp Berthold, Bianca Forkel et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.