Skip to content

Category

robotics

1,156 papers

#artificial intelligence Preprint Open access Oct 2026

Learning Visual Feature-Based World Models via Residual Latent Action

World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features instead of raw video pixels, offering a promising alternative that is more efficient and less prone to...

Xinyu Zhang, Zhengtong Xu, Yutian Tao et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Generalizable Dense Reward for Long-Horizon Robotic Tasks

Existing robotic foundation policies are trained primarily via large-scale imitation learning. While such models demonstrate strong capabilities, they often struggle with long-horizon tasks due to distribution shift and error accumulation. While reinforcement learning (RL) can finetune these models, it cannot work well...

Silong Yong, Stephen Sheng, Carl Qi et al. · 0 citations
#machine learning Preprint Open access Oct 2026

HARL-A: An Extensible Benchmark Framework for Heterogeneous Multi-Agent Adversarial Reinforcement Learning in IsaacLab

Progress in adversarial multi-agent reinforcement learning (MARL) for robotics has been hampered by a lack of shared, extensible infrastructure that supports heterogeneous agent morphologies in high-fidelity physics simulation. Existing frameworks either focus on cooperative tasks, rely on simplified physics engines, o...

Isaac Peterson, Christopher Allred, Jacob Morrey et al. · 0 citations
#machine learning Preprint Open access Oct 2026

QF3: Fast Flow RL with Filtered Q-Gradients

Flow policies have become a standard policy class for learning robot behaviors from demonstrations, but reinforcement learning is still critical for improving pre-trained flow policies or learning them from scratch through interaction. We introduce QF3 (Fast Flow RL with Filtered Q-Gradients), an online off-policy RL a...

Chung Min Kim, Brent Yi, David McAllister et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Compact Robot Policies Need Fine-Grained Visual Representations

Multi-task manipulation policies differ in architecture, scale, and pretrained priors all at once, so published comparisons cannot attribute performance to any single component. We argue that most of it comes from the visual representation, and that parameter scale and generative priors are largely incidental. To test...

Na Chen, Run-Qiu Yang, Jia-Wei Tang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Beyond Waypoint Regression: Query-Based Cost Learning over Reachable Ego Futures for End-to-End Driving

End-to-end planners based on waypoint regression achieve strong open-loop accuracy, but they primarily learn to mimic expert geometry and remain difficult to adapt to deployment-time safety constraints. We propose a query-based cost-learning framework that estimates bounded costs for dynamically reachable ego trajector...

Ahmed Abouelazm, Rupert Polley, Qingyuan Zhang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

Energy-Aware Path Following: Comparative Analysis of Reinforcement Learning and NMPC for Electric Vehicles

Path-following control strategies typically follow the bi-objective optimization dilemma: minimizing deviations from a reference path while maintaining smooth speed profiles. The latter objective is especially relevant for Electric Vehicles (EVs), since their limited driving range can be extended by recovering energy t...

Mohamed Sabaa, Mostafa Emam · 0 citations
#machine learning Preprint Open access Oct 2026

Learning Grasp Targeting from Point Clouds for Log Pile Clearing on a Hydraulic Crane

In mill yards, log loaders clear dense piles by a sequence of bundle grasps: hundreds of logs rest in contact, and each removal changes the pile available to the next grasp. A learned policy chooses where to place and orient the grapple from unsegmented point clouds and runs on a trailer-mounted hydraulic forestry cran...

George Sideris, Lucas Bessai, Heshan Fernando et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Modeling Latent Disturbances for Robust Decision-Making in World Models

In this paper, we study robust decision-making in the latent space of world models (WMs). Robust optimization is a mathematical framework where, given explicitly specified dynamics and physically meaningful disturbances, a robot can select actions that remain effective even under worst-case disturbances. However, apply...

Junwon Seo, Andrea Bajcsy · 0 citations
#machine learning Preprint Open access Oct 2026

The Robot Is Not Its Description: GaugeBench for Representation Robustness in Morphology-Aware Policies

A robot description does more than specify a physical mechanism: it also encodes arbitrary conventions, such as joint-axis direction, joint-angle zero, and the order and names of links and joints. Morphology-aware policies consume interfaces built from these descriptions, yet cross-embodiment evaluation typically chang...

Rahath Malladi, Arshia Sangwan, Rajesh K. Gupta et al. · 0 citations
#machine learning Preprint Open access Oct 2026

BiGym 2.0: Benchmarking Learned and Agent-Developed Policies for Humanoid Household Manipulation

Humanoid household manipulation requires the arms to act while the body balances, steps and changes posture. We present BiGym 2.0, an adaptation of BiGym for the Unitree G1 across 20 household tasks using a unified whole-body controller for demonstration and evaluation. The suite provides 60 native human virtual-realit...

Zexi Zhang, Zecheng Zhu, Zidong Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Seeing the Invisible: Physics-Guided Visual Prompting for Temperature- and Radiation-Aware VLA Navigation

Vision-Language-Action (VLA) models have become a major paradigm for Vision-and-Language Navigation (VLN). However, in safety-critical facilities, invisible risks such as radiation or temperature spikes cannot be detected by an RGB camera, and handling each risk is expensive, requiring a new encoder, new data, and mode...

Hojoon Son, Fan Zhang · 0 citations

From tech blogs

See all →
Microsoft Research Blog Sep 23, 2026

Offloaded inference for real-world physical AI robotics

Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.