Jul 2026· 2026 23rd International Conference on Ubiquitous Robots (UR)· pp. 390-396· 0 citations· 17 references
Computer Science
Abstract
Robot skill generation is often approached from two distinct perspectives: normative trajectory optimization, which emphasizes smoothness-based criteria such as minimum jerk, and imitation-based learning, which prioritizes fidelity to demonstrated behaviors. While both paradigms aim to produce feasible and meaningful motions, they are typically formulated separately. In practice, however, many robotic skills require trajectories that are both dynamically smooth and faithful to demonstrations. We propose a unified framework for normative–imitative trajectory optimization that makes this trade-off explicit and tunable. Our framework formulates trajectory generation as a constrained quadratic program combining weighted linear-operator smoothness penalties, a quadratic imitation anchoring term, and affine equality constraints for feasibility. For affine-constrained instances, the resulting problem is strictly convex, admits a unique global minimizer, and can be solved efficiently using standard quadratic programming techniques. Our proposed formulation unifies a broad class of smoothness objectives, including minimum velocity, acceleration, jerk, snap, and elastic energy models, within a single operator-based representation, while incorporating demonstration fidelity in a principled manner. Simulation and real-world experiments on a UR5e robotic arm demonstrate that our framework provides predictable interpolation between purely normative and purely imitative behaviors, offering a compact and extensible foundation for trajectory learning from demonstration.
This report adopts a two-piece MINCO parameterization, trading time for smoothness without altering the trajectory's spatial profile, and replaces score regression with a ranking loss, preventing small score errors from reordering the candidate set.
Dual-arm manipulation or physical human-robot coordination requires robots to adapt rapidly to changing environments and constraints. Traditional Learning from Demonstration approaches struggle to generalize when faced with out-of-distribution scenarios, requiring costly retraining. We propose a Movement Primitive learning algorithm based on Gaussian Processes, combined with real-time zero-shot adaptation through Pathwise Conditioning. The method encapsulates the predictive uncertainty of the demonstrated movement using heteroscedastic GPs and utilizes an update via Matheron's rule to instantaneously adjust the trajectory to new via-points, without the need to retrain the underlying model. This formulation is extended to dual-arm coordination by dynamically calculating 6D relative constraints to maintain a closed kinematic chain. Experimental results, both in 2D comparisons against task-parameterized models and in tasks with the ADAM robot, demonstrate robust adaptation with near-zero error in real time, making it applicable for highly changing environments.
Adrián Prados, L. Lishan, Alberto Mendez et al.· Jornadas de Automática· 0 citations
Dynamic Movement Primitives (DMPs) are widely used for robot skill reproduction from demonstrations, but pose trajectory reproduction for continuous manipulation tasks remains challenging because translational and rotational motions must be represented consistently while maintaining accuracy, disturbance recovery, and terminal smoothness. Existing screw-displacement pose DMPs provide a geometrically consistent formulation on SE(3); however, their isotropic fixed-gain feedback limits direction-dependent correction, and deterministic forcing terms do not provide an explicit estimate of prediction reliability, which may cause over-shaping and high terminal jerk. This paper proposes a direction-adaptive and uncertainty-weighted pose DMP framework for robot skill reproduction from pose trajectories obtained from demonstrations. A Riemannian Motion Policy (RMP)-inspired direction-adaptive feedback mechanism is introduced to adjust recovery and damping according to the current pose error directions and magnitudes, improving trajectory-level correction and disturbance recovery. In addition, a Sparse Spectrum Gaussian Process (SSGP) is used to model the forcing term probabilistically, and its predictive variance is combined with a phase-dependent gate to attenuate low-confidence forcing contributions, particularly near the terminal phase. Simulation studies on RoboMimic trajectories show that the RMP-inspired feedback primarily improves pose reproduction accuracy and disturbance recovery, whereas the SSGP-based weighting substantially reduces terminal translational and rotational jerk, with a slight accuracy compromise relative to RMP-DMP. A papermaking robot case study further demonstrates the deployment feasibility of the generated pose trajectories on a real continuous-operation platform.
In this article, we propose trajectory-constrained human-guided reinforcement learning (TCHug-RL), a new framework that 1) produces guidance at the full-trajectory level to suit the planning horizon of autonomous driving and 2) provides a formal criterion for deciding when and how to inject human input into the policy-update loop. First, we propose a trajectory-level similarity of human guidance, which is defined such that the satisfaction of trajectory-level similarity implies the satisfaction of single-step similarity. Consequently, trajectory-level similarity is strictly stronger than the single-step one. Then, the trajectory-level similarity is modeled as a constraint to guide the subsequent policy optimization. In this way, not only is a clear criterion for human guidance provided, but the influence of such guidance on the policy is also transparent and interpretable. Finally, we leverage the Lagrangian method to solve the constrained optimization problem and provide a strong duality analysis for it in the unparameterized policy space. Furthermore, a practical algorithm implementation of TCHug-RL and experiments are provided, demonstrating that TCHug-RL effectively leverages human guidance and achieves improvements of 22.1% in learning efficiency and 18.7% in overall performance in autonomous driving tasks, compared to the state-of-the-art methods.
Li-Fei Dai, Yiqun Liu, Hao Zhang et al.· IEEE Transactions on Neural...· 0 citations
This work replaces the costly convex relaxation step required by nominal GCS with a single forward pass through a Graph Attention Network that predicts a set of highly probable candidate paths through the graph, and generates a lightweight ranking network that orders these candidates by their estimated trajectory cost.
Ananya Trivedi, Sarvesh Prajapati, M. K. M. Jaffar et al.· 0 citations
Autonomous driving in complex urban environments requires trajectory planning that balances safety, efficiency, and human-like behavior. Although imitation learning (IL) can capture expert driving patterns from large-scale demonstrations, existing IL-based planners still face challenges in safety-critical scenarios and long-tail traffic distributions. Meanwhile, optimization-based planners provide explicit constraint handling but are often separated from upstream learning modules, limiting their ability to jointly improve trajectory generation and planning feasibility. To address these issues, we propose a hybrid trajectory planning framework that integrates IL-based multimodal trajectory proposal with differentiable optimization. In the proposed framework, an IL backbone generates candidate ego trajectories and surrounding-agent predictions, while a differentiable optimizer refines the selected trajectory using multi-objective cost functions with learnable weights related to safety, efficiency, and comfort. This design enables optimization objectives and constraints to provide gradient feedback to the upstream planning network, improving the consistency between candidate generation and downstream planning objectives. In addition, we introduce a surrounding agent centric data augmentation strategy that reuses real-world trajectories of surrounding vehicles as additional expert demonstrations, thereby enriching complex interaction and long-tail scenarios without extra data collection. Closed-loop experiments on the nuPlan benchmark show that the proposed method achieves a composite score of 94.04, outperforming PLUTO’s 93.14 while using only 30% of the training data. The results demonstrate that the proposed framework improves closed-loop planning performance, trajectory feasibility, and data efficiency under complex urban driving scenarios.
Shihao Zhang, Ziyu Song, Zhaochen Xia et al.· Proceedings of the Instituti...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.