Residual Deep Reinforcement Learning-Based Computed Torque Control for a Cable-Driven Lower-Limb Rehabilitation Robot under Disturbances and Parametric Uncertainties
Bounded residual learning is supported as a practical robustness-enhancement strategy for simulation-based rehabilitation robot control and will motivate further constraint-aware and experimental validation.
Abstract
Accurate trajectory tracking in cable-driven lower-limb rehabilitation robots is challenging because model uncertainty, external disturbances, joint constraints, and pull-only cable actuation can degrade nominal control performance. Conventional model-based controllers provide an interpretable control structure but remain sensitive to model mismatch, whereas fully learning-based control can reduce transparency and complicate constraint-aware operation. This study proposes a residual deep reinforcement learning-enhanced computed torque control framework in which computed torque control generates the nominal command and a bounded Deep Deterministic Policy Gradient policy supplies only an additional compensating torque. The approach is evaluated in simulation under nominal, uncertain, disturbed, combined, and generalization conditions, together with trajectory-tracking, joint-limit, cable-demand, workspace-feasibility, and cable-Jacobian diagnostics. Across the evaluated conditions, the residual controller improves tracking and disturbance rejection relative to computed torque control while preserving the interpretable model-based command structure and satisfying the reported feasibility checks in the representative evaluation. Broader tests indicate that tracking improvements can persist beyond the representative case while also exposing trajectory-dependent constraint limitations. These results support bounded residual learning as a practical robustness-enhancement strategy for simulation-based rehabilitation robot control and motivate further constraint-aware and experimental validation.
This thesis investigates advanced modeling and control strategies for robotic manipulators, focusing on the DLR-HIT II robotic hand and the KUKA LBR iiwa. It presents three core contributions that integrate simulation, model-based control, and data-driven methods to improve torque and position control under uncertainties and disturbances. First, a dual-platform simulation framework is developed using MATLAB Simscape Multibody and CoppeliaSim. The system accurately models the DLR-HIT II hand’s kinematics and dynamics, enabling both control validation and realistic interaction with virtual environments. The use of Unified Robot Description Format (URDF)-based modeling sup-ports reusability and modular analysis. Second, a physics-informed neural network (PINN) is proposed for direct torque and position control. This method uses only time and joint position inputs, internally computes derivatives, and generalizes well across various trajectory types. It eliminates the need for separate feedback controllers and shows strong robustness under disturbances, while maintaining low computational cost. Third, a Spike-Aware Hybrid Torque Control (SA-HTC) architecture is enhanced with a feedforward neural network (NN) trained offline. The network learns to improve Computed Torque Control (CTC) for unmodeled effects, friction, and external forces by refining the CTC torque output in real time. Simulation results across diverse trajectories and noise levels demonstrate that the SA-HTC method significantly improves tracking accuracy and robustness compared to classical CTC. To support deployment on position-controlled hardware, a physics-informed torque-to-position interface is introduced and compared with virtual stiffness and admittance wrappers, yielding lower steady-state bias, reduced phase lag, lower tracking error, and robust cross-trajectory generalization. Thus, these contributions advance the integration of learning-based and model-based control strategies in robotics. The results highlight scalable and efficient methods for accurate trajectory tracking and torque control, offering practical potential for robotic manipulation and automation applications.
This paper presents an advanced deep reinforcement learning (DRL) framework for precise trajectory tracking control of an underactuated 2-degree-of-freedom (2-DOF) helicopter system using the twin delayed deep deterministic policy gradient (TD3) algorithm. The 2-DOF helicopter serves as a benchmark for nonlinear, coupled, and underactuated systems, posing significant challenges for conventional control approaches. Both classical linear and nonlinear control methods provide baseline solutions; however, their performance often degrades in the presence of parameter variations, uncertainties, and external disturbances. To overcome the severe value overestimation errors caused by aerodynamic cross-coupling in standard actor-critic architectures, a model-free TD3-based controller is developed, incorporating an artificial potential field-inspired reward function to simultaneously optimize tracking accuracy, energy efficiency, and control smoothness. Compared with standard DRL approaches such as the deep deterministic policy gradient (DDPG), the TD3 algorithm addresses key limitations by employing twin critics to reduce overestimation bias, delayed policy updates to improve training stability, and target policy smoothing to enhance robustness. Comprehensive simulations conducted in a MATLAB/Simulink environment demonstrate the superior performance of the proposed TD3 controller compared to classical and intelligent approaches, including proportional-integral-derivative (PID), fuzzy PD + I, and fuzzy PD + FF controllers. For multi-step trajectory tracking, TD3 reduces overshoot to 4.6% (pitch) and 3.8% (yaw) compared to 22.4% and 18.7% for PID, while decreasing settling time by up to 66%. The steady-state error is reduced to 0.18° (pitch) and 0.15° (yaw), representing improvements exceeding 80% over PID. In addition, TD3 minimizes cross-coupling effects by over 60%, enabling effective decoupled control of pitch and yaw dynamics. Under complex trajectories and disturbance conditions, including ± 10% parametric uncertainties and external torque disturbances, the TD3 controller consistently achieves the lowest tracking errors, fastest convergence, and smoothest control signals, reducing control variation by up to 65% compared to conventional methods. These results highlight the effectiveness of TD3 for controlling nonlinear and underactuated systems and provide a solid foundation for future experimental validation and real-world deployment in aerial robotic platforms operating in uncertain environments.
Zied Ben Hazem, Muhammed Özdemir, Firas Saidi et al.· Discover Robotics· 1 citation
Bilateral teleoperation systems that include joint flexibility better reflect real robotic systems used in surgery, space, and rehabilitation. However, joint flexibility together with time-varying communication delays makes it difficult to maintain stable and coordinated motion between the master and slave robots. To address this, we propose a hybrid control method that combines a stable Proportional-plus-Damping (P+d) controller with a model-free deep reinforcement learning agent based on the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm. The P+d controller provides basic stability under bounded delays, while the learning agent adjusts and tunes the remote-side proportional and damping gains in real time to reduce vibrations and improve tracking. Stability is guaranteed for bounded time-varying delays using Lyapunov-Krasovskii analysis. The approach provides a practical solution for teleoperation systems facing both joint flexibility and uncertain network delays.
Armin Attarzadeh, Mohammadali Ghaemifar, A. Khanzadeh et al.· arXiv.org· 2 citations
This study addresses the trajectory tracking control problem for an underactuated hovercraft subject to additive bias and multiplicative loss-of-effectiveness thruster faults under environmental disturbances. In these systems, actuator degradation structurally breaks the differential flatness mapping, driving nominal controllers to generate control actions that induce severe actuator saturation and cause instability. To resolve this challenge, a hierarchical physics-informed neural adaptive control (PINAC) framework is proposed. First, a gated-recurrent-unit physics-informed neural observer (PINO) is designed to isolate thruster faults from exogenous hydrodynamic disturbances. Second, a constrained Safe-TD3 reinforcement learning agent functions as a supervisor, computing an online dilation factor to slow down the mission timeline, thereby reconfiguring the reference trajectory to accommodate degraded actuator boundaries. Third, a low-level non-singular terminal sliding mode (NTSM) controller is implemented as a tracking-guarantee layer. Unlike classical asymptotic schemes where convergence is only achieved as time approaches infinity, or finite-time controllers where the settling time depends on the initial state, the proposed PINAC framework guarantees practical fixed-time stability, ensuring that the settling-time bound is independent of initial conditions. Simulation results demonstrate that the designed controller prevents actuator saturation, provides smooth trajectory adjustment, and reduces tracking errors under severe composite faults.
Shafqat Ali, Aamir Mehmood, Faiza Iftikhar et al.· Journal of Marine Science an...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.