Aug 2026· Frontiers in Robotics and AI· Vol 13· 0 citations· 55 references
Medicine
TL;DR
This study investigates the combination of a state-of-the-art reinforcement learning (RL) algorithm with human demonstrations to learn how to open a door with minimal task-specific engineering on an articulated soft robot arm and shows that combining LfD with RL results in both better performance and more robust behaviors.
Abstract
Learning from demonstration (LfD) has become a popular approach with the emergence of modern transformer-based algorithms. However, the performance of these policies is limited by the quality of the demonstrations. Combining imitation and exploration promises to train policies that perform better and are more reliable. However, this requires a robotic system that can explore safely without damaging itself or the environment, especially in contact-rich tasks during which the robot must exert force on its environment to solve the task. In this study, we investigate the combination of a state-of-the-art reinforcement learning (RL) algorithm with human demonstrations to learn how to open a door with minimal task-specific engineering on an articulated soft robot arm. We found that learning from both exploration and demonstration data stored in separate buffers makes the algorithm not only more sample-efficient and robust but also allows the policy to reach a higher performance level than the provided expert demonstrations. We also show that using an articulated soft robotic arm allows us to perform RL on a real robotic system without any pretraining and with a simple safety system that does not require any additional sensors, such as force–torque sensors. Additionally, we can implicitly learn the nonlinearities stemming from the soft materials in the actuator. Our findings show that combining LfD with RL results in both better performance and more robust behaviors and indicate that articulated soft robots allow for learning contact-rich tasks safely on a real system.
This study proposes a new method, Wrench-Augmented Reinforcement Learning (WARL), which introduces a wrenche (force and torque) into the action space, and shows that introducing a wrench can encourage behaviors that do not sufficiently exploit the robot's physical embodiment.
A unified framework that combines centralized training with decentralized execution (CTDE) and a Hybrid Reward Architecture (HRA) is introduced that enables multiple actors to share a centralized multi-head critic and substantially improves both sample efficiency and policy performance.
Changhao Li, Yifang Zhang, Heng Zhang et al.· 0 citations
Sample effective and stable training remains a key challenge in reinforcement learning (RL), especially for real-world applications such as mobile robot control where data collection is time-consuming and failures may be hazardous.Building on the residual reinforcement learning paradigm, this work presents, to the best of our knowledge, one of the first detailed physical studies of a residual Soft Actor-Critic (SAC) controller for camera-based lane following on a mobile robot. We combine an established stable, but sub-optimal lateral P-controller with a regularized SAC agent in a hybrid architecture. The classical controller provides baseline stability and rapid initial learning, while the RL agent learns residual corrections to improve performance. We employ a PID-inspired reward function and quadratic policy output regularization to ensure smooth control actions and effective sim-to-real transfer.The hybrid controller design enables rapid training convergence, requiring only a few epochs and outperforming the pure RL approach by two orders of magnitude in sample efficiency. This enables efficient hyperparameter tuning in simulation and opens the door to future learning directly on physical robots. Fine-tuning with only a few dozen real-world laps achieved robust transfer to the physical robot, maintaining the same architecture and hyperparameters. The method generalized effectively to new scenarios, such as lane changes.
Fedi Boukhris, J. Will, Timo von Marcard et al.· International Conference on...· 0 citations
Learning from Demonstration (LfD) allows robots to learn manipulation tasks directly from humans, thereby supporting the versatile application of robots. Most LfD methods do not explicitly model the physical interactions between a robot and its environment, such as the making and breaking of contact, while these are crucial during manipulation tasks. Because the same basic physical interactions recur often, they can be a basis for robust, generalizable, and adaptive task reproduction. We propose an LfD method that explicitly uses what physical interactions take place where and when. Using that information, a hybrid position-force controller tracks demonstrated trajectories until contact-based transition conditions from the demonstrations are met. We evaluate our method in real robot experiments consisting of opening doors and locks, bolt picking and screwing, dislodging, and surface contouring. We show that explicitly modeling physical interactions benefits LfD in four ways. First, by allowing reproduction of complex, sequential, and contact-rich manipulation tasks using only a single demonstration and no prior knowledge of the task. Second, by facilitating robustness to unknown geometric variations in the environment. Third, by facilitating generalization when geometric variations are known. Fourth, by facilitating online adaptation using geometric information explored during task reproduction. We discuss how robustness, generalization, and adaptivity can be explicitly implemented, which is generally lacking in the LfD literature. Thereby, our work aims to close a gap in interpretable few-shot LfD of robotic manipulation.
A. H. G. Overbeek, H. V. D. Kooij, M. Vlutters· 0 citations
This work uses Sample-based Model Predictive Control entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets and validate the robustness of this sim-to-real framework by successfully deploying complex loco-manipulation skills across different morphologies.
Martin Schuck, Maks Sorokin, S. Manni et al.· 0 citations
It has always been expected that robots can actively manipulate complex environments to fulfill human requirements. This process typically necessitates that the robot be equipped with the ability for embodied exploration and manipulation. To achieve this goal, in this paper, we propose to incorporate multi-source knowledge to enhance the ability of robotic embodied exploration and manipulation. Specifically, to eliminate the inherent biases in decision-making of large language models (LLMs), we introduce a multi-source knowledge fusion module to generate more reasonable exploration sequences. Notably, grasping detection plays a critical role in the process of robot manipulation. To achieve a better balance between the accuracy and efficiency of the grasping detection network, we design a two-branch feature fusion module with residual blocks to improve network performance. Conditioned on the aforementioned innovations, the robot is capable of actively exploring and manipulating in constrained environments to meet human requirements. Extensive experiments are conducted in both simulation and real-world environments. The results demonstrate the effectiveness and efficiency of our proposed framework.
Jin Liu, Kai Sun, Leibing Xiao et al.· Robotica (Cambridge. Print)· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.