This work introduces a novel inductive bias towards simple policies in reinforcement learning by minimizing the entropy of entire action trajectories, corresponding to the number of bits required to describe information in action trajectories after the agent observes state trajectories.
Bang You, Chenxu Wang, Wen-Ju Yang et al.· 0 citations
Learning task-relevant representations is crucial for reinforcement learning. Recent approaches aim to learn such representations by improving the temporal consistency in the observed transitions. However, they only consider individual transitions and can fail to achieve long-term consistency. Instead, we argue that ca...
Bang You, Huaping Liu, Jan Peters et al.· 0 citations
A pretrained robot foundation policy may execute most of a long-horizon task yet repeatedly fail at a few critical subtasks. Collecting additional full-task demonstrations for supervised fine-tuning (SFT) requires operators to repeat behaviors the policy already performs well. Reinforcement learning (RL) fine-tuning of...
Sichang Su, Benjamin Yang, Zhiyun Deng et al.· 0 citations
The proposed epistemic uncertainty-driven adaptive rollout strategy for offline world model training following an auto-curriculum training scheme indicates that epistemic uncertainty is useful not only for downstream policy regularization, but also for making world model training itself more compute-efficient.
Nikodem Sebastian Zymla, Laurin Thiele, Johannes Pitz· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
FootQuery is presented, a perceptive locomotion framework that queries depth history using each foot's predicted next touchdown using each foot's predicted next touchdown to organize visual history around anticipated contacts for perceptive humanoid locomotion.
Octopus crawling motivates soft robots that exploit redundancy, yet discovering and organizing diverse coordination modes for adaptation remains challenging. To address this, we introduce a Diffusion-based Uncertainty-aware Optimization (DUO) algorithm that learns demonstration-free crawling controllers for a simulated...
Seung Hyun Kim, Heng-Sheng Chang, Kimia Kazemi et al.· 0 citations
This work designs each task curriculum with compositional tasks that combine aspects of the tasks seen in the sequence, and factorise this composition along the axes of action and perception to better understand how different input modalities bottleneck knowledge reuse.
Hao-Yu Zhou, Joe Watson, Anson Lei et al.· 0 citations
A systematic comparison of two state-of-the-art motion-imitation reinforcement learning (MIRL) pipelines, one built on SCONE/HyFyDy and one built on MuJoCo/MyoSim, concludes that the more advanced physiological realism of HyFyDy currently makes it more suitable for musculoskeletal modeling.
Ayah Ahmad, Claire E. Borden, Maegan Tucker· 0 citations
The dominant method of processing sonar data is using image-based representations, requiring the preprocessing of image data on autonomous systems. We propose an alternative data processing method for remote sensing applications via the use of data in Comma-Seperated Value format. Experimentation on our alternative app...
Logan Luna, Sirio Jansen-S\'anchez, Ilteris Demirkiran et al.· 0 citations
Recent advances in large language models have improved their effectiveness as back-end components for voice assistants, particularly in intent understanding and context-aware input classification. However, online-hosted models introduce network dependency and variable inference latency, limiting their suitability for t...
Daniel Henel, Frederik Werner, Alexander Langmann et al.· 0 citations
Reinforcement learning (RL) controllers have been recently adopted for Unmanned Aerial Vehicles (UAV) navigation and control. However, they are susceptible to action-space attacks that overwrite the action commands after the policy generates them and before the actuators execute them. While most existing defenses targe...
Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that may limit interpretability of learned policies and may lack explicit control over gait behaviors. We introduce a framework where distinct gaits are specified using parameterized constraints...
Merve Atasever, Keyan Azbijari, Cagan Bakirci et al.· 0 citations
Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads. The post Offloaded inference for real-world physical AI robotics appeared first on Microsoft Research.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.