Robot Trajectron V3 (RT-V3), a probabilistic shared control framework designed for grasping tasks, formulates shared control as Bayesian inference by learning a prior over user intent and combining it with real-time user commands to estimate the posterior intent distribution.
Abstract
We aim to address the challenge of teleoperating robotic arms for high-degree-of-freedom (high-DoF) manipulation tasks, which is cognitively demanding and error-prone, particularly when relying on low-bandwidth interfaces. We propose Robot Trajectron V3 (RT-V3), a probabilistic shared control framework designed for $SE(3)$ grasping tasks. RT-V3 formulates shared control as Bayesian inference by learning a prior over user intent and combining it with real-time user commands to estimate the posterior intent distribution. The prior models user intent as a distribution over future trajectories conditioned on past robot dynamics and visual scene context. The intent prior is parameterized by a transformer-based conditional generative model that reasons over point clouds and candidate grasp poses, together with a factorized translation-rotation representation that improves learning efficiency in high-dimensional action spaces. During execution, RT-V3 continuously estimates the posterior distribution over future trajectories by combining the learned intent prior with a user-command likelihood derived from the observed control input, enabling continuous intent refinement and shared assistance. Comprehensive experiments demonstrate that RT-V3 achieves high accuracy in trajectory prediction and competitive performance in reactive planning. Furthermore, real-world user studies indicate that RT-V3 significantly outperforms baseline methods in terms of success rate and efficiency, while substantially reducing the user's physical and mental workload.
GeniWorld is presented, an interactive world model for robots that generalizes robustly across unseen scenarios by explicitly decoupling embodiment kinematics from environmental dynamics, and generates diverse manipulation trajectories within the world model, improving downstream policy performance and robustness in complex environments.
An integrated framework that combines a vision module for landing point and time prediction with a direct position and time conditioned RL locomotion policy, instead of intermediate velocity commands is proposed, which mitigates perception latency during dynamic interception.
Yi-Dong Zhu, Zibo Dai, Tong-Ning Zhang et al.· 0 citations
This paper proposes a nested kino-dynamic framework for rapid feasibility checking and dynamically consistent trajectory generation given a candidate contact sequence and shows that the generated trajectories can be tracked using a reinforcement learning (RL)-based controller and are of sufficiently high quality for execution in real-world loco-manipulation scenarios.
Michal Ciebielski, Shafeef Omar, Aaron M. Johnson et al.· arXiv.org· 0 citations
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in generalized robotic control, yet their scalability is fundamentally bottlenecked by the high cost and low diversity of teleoperated data. While abundant, human demonstration videos cannot be directly utilized for policy training due to the severe morphological differences between human anatomy and robotic manipulators. To bridge this embodiment gap, this work proposes a lightweight retargeting pipeline that kinematically retargets human interaction data (DexYCB) onto six-degree-of-freedom manipulator trajectories to fine-tune policies based on the pi0.5 architecture. By prioritizing Cartesian positional alignment via constrained Inverse Kinematics (IK) and introducing an object-based grasping heuristic, smooth geometric priors are generated without relying on computationally heavy visual synthesis. Physical evaluations demonstrate that retargeted models significantly outperform standard teleoperation (40.6% success rate), achieving 65.6% success via co-training and a peak 78.1% success rate via two-stage cross-embodiment co-training. Furthermore, evaluations under extreme visual clutter reveal that explicitly retargeted policies exhibit immunity to semantic visual distractors. Finally, we examined and analysed Terminal State Ambiguity, a temporal failure mode where generative models fail to terminate the task when exposed to scenarios similar to the nature of human video priors.
José António Nunes Andrade, Mauro Castelli· Decision Making Advances· 0 citations
Deep reinforcement learning is emerging as a powerful alternative to traditional inverse kinematics for controlling robotic manipulators. By learning optimal actions through interaction with the environment, it enables adaptable and precise control in complex continuous workspaces, making it suitable for dynamic manipulator operations. This paper proposes an actor–critic deep reinforcement learning framework for manipulator control in a continuous workspace. Conventional inverse kinematics solutions can become computationally complex for high-degree-of-freedom manipulators or complex kinematic structures. Accurately generating joint angles by deriving configuration-specific equations therefore remains a persistent challenge. This study employs deep deterministic policy gradient and twin delayed deep deterministic policy gradient algorithms to directly predict joint angles for target-reaching tasks, eliminating reliance on conventional inverse kinematics formulations. A key contribution is a simple yet effective linear reward function that supports stable convergence in continuous action spaces. The proposed framework is implemented using ROS2 and Gazebo simulation and validated on a custom-built physical manipulator. The work addresses complex equation-based control through reinforcement learning and demonstrates proof of concept on a 3-degree-of-freedom manipulator. Experimental results show high task success rates of 95.77% for the deep deterministic policy gradient algorithm and 94.71% for the twin delayed deep deterministic policy gradient algorithm, with accurate end-effector positioning within 0.01 m. These results confirm the effectiveness of the proposed framework for accurate and reliable manipulator control with reduced model complexity and improved real-world applicability.
Received: 13 September 2025 | Revised: 14 April 2026 | Accepted: 10 June 2026
Conflicts of Interest
The authors declare that they have no conflicts of interest to this work.
Data Availability Statement
Data are available from the corresponding author upon reasonable request.
Author Contribution Statement
Shivkumar Sankaralingam: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Writing – original draft, Writing – review & editing, Visualization. NippunKumaar Arulmani Angamuthu: Conceptualization, Methodology, Formal analysis, Investigation, Writing – review & editing, Visualization, Supervision. Suja Palaniswamy: Writing – review & editing, Visualization, Project administration, Funding acquisition.
A neural reduced dynamics framework is developed that separates the state the model propagates from what can be supplied as an input or recovered analytically, trains policies entirely inside the frozen learned model, and validates them back in the high-fidelity simulator.
Harry Zhang, Dan Negrut· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.