2026· Computers, Materials & Continua· Vol 88, pp. 1-10· 0 citations· 66 references
TL;DR
This comprehensive review systematically synthesises state-of-the-art algorithmic building blocks across perception, dynamics modelling, and control to ensure that the next generation of embodied AI achieves human-level intelligence while strictly meeting the safety, accountability, and regulatory requirements for dependable clinical and industrial deployment.
Abstract
: The integration of Deep Learning, Deep Reinforcement Learning, and massive Vision-Language-Action (VLA) foundation models has catalysed a profound paradigm shift in robotics, transitioning systems from rigid automation to dynamic, open-world autonomy. Despite transformative breakthroughs in fields such as healthcare, ranging from adaptive robotic rehabilitation to autonomous surgical manipulation and silver care, widespread real-world deployment remains severely bottlenecked. This limitation primarily stems from the “Reality Gap” inherent to sim-to-real transfer and a fundamental epistemological tension: the stochastic, “black-box” nature of unconstrained neural networks fundamentally conflicts with the deterministic, zero-violation safety guarantees demanded by physical robotics. To address these critical barriers, this comprehensive review systematically synthesises state-of-the-art algorithmic building blocks across perception, dynamics modelling, and control. Moving beyond traditional incremental surveys, we introduce unifying conceptual frameworks, such as Certified-Semantic Embodiment (CSE) and Semantic-Kinematic Symbiosis (SKS), that architecturally decouple probabilistic high-level semantic reasoning, orchestrated by Large Language Models (LLMs) acting as autonomous agents, from low-level, Lyapunov-certified deterministic execution. Furthermore, we formalise the evaluation pipeline for deployment realities, recommending a shift from empirical success rates to mathematically bounded frameworks such as Prediction-Powered Inference (PPI) to ensure robust sim-to-real generalisation. Ultimately, this review provides a rigorous technical roadmap for bridging the semantic-kinematic divide. By integrating cognitive adaptability with rigorous physical constraints, we aim to ensure that the next generation of embodied AI achieves human-level intelligence while strictly meeting the safety, accountability, and regulatory requirements for dependable clinical and industrial deployment.
This short course presents a unified pipeline for developing humanoid and general-purpose robot policies, spanning synthetic data generation, policy training, and deployment, and gains a practical understanding of how simulation, world models, and foundation models compose into a scalable, end-to-end system for generalizable physical AI.
Edith Llontop, A. Santhosh· Proceedings of the Special I...· 0 citations
Reinforcement learning (RL) has emerged as a promising approach for autonomous surgical robotic subtasks. Recent advances include deep reinforcement learning (DRL), imitation learning (IL), and vision–language–action (VLA) models. However, current evidence remains fragmented across simulation benchmarks, task-specific demonstrations, and limited clinical studies. Existing reviews primarily focus on RL algorithms, while the broader pathway from algorithm development to clinically deployable surgical autonomy has not been comprehensively synthesised. This PRISMA 2020-guided systematic review examines RL, IL, safe RL, simulation-to-real (sim-to-real) transfer, foundation models, VLA systems, and regulatory readiness in surgical robotics. We searched IEEE Xplore, PubMed/MEDLINE, Embase, Scopus, Web of Science, the Cochrane Library, ACM Digital Library, arXiv, and medRxiv for studies published between January 2015 and March 2026, with additional studies identified through backward citation tracing. Eligible studies proposed novel RL, imitation learning, or foundation-model approaches for surgical robotics with empirical validation in simulation or on physical robotic platforms. Two reviewers independently extracted data using a predefined coding scheme, and a third reviewer resolved disagreements. Owing to substantial heterogeneity in platforms, tasks, and outcome measures, a quantitative meta-analysis was not feasible; therefore, the evidence was synthesised narratively using a comparative framework. A total of 220 studies met the inclusion criteria, covering eleven active surgical RL platforms, seven paired sim-to-real studies, emerging foundation-model architectures, and three FDA-cleared robotic systems exhibiting Level 3 autonomy. Available comparative studies suggest that hierarchical approaches can outperform flat policies in long-horizon tasks, while language-conditioned models demonstrated promising multi-step surgical capabilities. Seven paired simulation-to-real studies were identified, encompassing tissue retraction, guidewire navigation, and surgical cutting tasks. Sim-to-real performance gaps varied substantially by task and metric, with success-rate gaps ranging from −10 to 50 percentage points (negative values indicating better real-world than simulated performance), while paired mean spatial errors differed by at most 0.61 mm. Most studies employed domain randomization or visual domain adaptation; hierarchical reinforcement learning demonstrated advantages over flat policies in multi-step surgical tasks. Explicit safety-constrained methods (CPO, CBF, and SER), formal verification, and regulatory-aligned evaluation were reported in fewer than 3% of applied studies. Most evidence remained simulation-based, with no reported autonomous RL execution in vivo in humans. Overall, RL-based surgical robotics appears mature at the simulation stage but remains preclinical for autonomous clinical deployment. Future progress requires stronger sim-to-real validation, multimodal safety-aware architectures, alignment with IEC 62304, ISO 14971, FDA guidance, and the EU AI Act, and open benchmarks that jointly evaluate performance, safety, and surgeon trust.
M. Shahid, Abdullah, Zulaikha Fatima et al.· Biomimetics· 0 citations
Living organisms exhibit extraordinary adaptability to unseen environments through their intrinsic physical structures and lifelong feedback-driven learning. Endowing robots with comparable generalization is critical for reliable operation in the real world. While recent approaches attempt to improve generalization by scaling training data, such strategies remain impractical for robotics, where collecting real-world demonstrations at the scale of large language models is prohibitively costly and slow. Contrary to this reliance on massive datasets, we show that robots can generalize effectively under dynamics uncertainties even with limited training data by leveraging a feedback mechanism, namely PhyFilter, that corrects learning outputs with physics-filtered learning residuals. PhyFilter operates as a lightweight, model-agnostic module whose parameters can be automatically optimized through an auto-learning algorithm, eliminating manual tuning and enabling seamless integration with diverse robot policies. We validate PhyFilter across four representative robotic systems, demonstrating that it enables quadruped robots to generalize to unseen terrains, payload variations, and speed ranges; drones to flight under unseen wind disturbances; aerial manipulators to achieve centimeter-level in-air capture despite wind and mass uncertainties; and acceleration differentiators to remain robust with distribution shift. These results show that physics-filtered feedback can serve as a powerful alternative to massive data scaling.
Jindou Jia, Shixu Han, Meng Wang et al.· npj Robotics· 1 citation
Robotic systems are deeply embedded in both industry and everyday life, where they are expected to act with speed, precision, and reliability. Classical control and planning methods have long delivered strong guarantees, but often at the cost of computational efficiency and adaptability. More recently, learning-based approaches have shown promise in overcoming these limitations, enabling agents to leverage experience to accelerate decision-making and address previously intractable problems. In this work, we bridge these two approaches through a neuro-symbolic perspective on nonlinear motion planning. Inspired by the Thinking Fast and Slow paradigm, we introduce a dual-process architecture that combines the strengths of robust reasoning and learning. Our framework integrates state-of-the-art symbolic solvers as a ``System-2''component with experience-driven ``System-1''modules. A metacognitive controller dynamically orchestrates their interaction, selecting when to rely on fast intuition versus slower, more precise reasoning. By evaluating the framework across diverse nonlinear benchmark environments, we demonstrate that this architecture yields consistent gains in planning efficiency, accuracy, and generalization, while promoting reuse across tasks. The results suggest that tightly coupling learning with structured reasoning offers a scalable path toward more capable and adaptive robotic systems.
Jia-Yi Yan, Francesco Fabiano, Alessandro Abate· 0 citations
The Universal Manipulation Interface (UMI), originally developed by the Robotics and Embodied AI Lab at Stanford University, has demonstrated remarkable effectiveness for training manipulation policies for terrestrial robotic manipulators using imitation and diffusion based learning techniques. The long term objective of our research is to extend this technology to space robotic applications. There are numerous challenges largely unexplored, including harsh environmental conditions, stringent power and computational constraints, communication latency, and very limited opportunities for data collection and validation. This paper presents the first step toward achieving that goal by designing a similar gripper and testing it in both simulation and experiment with a Franka Emika robotic arm in our lab setting. We reproduce the data collection and diffusion policy training pipeline on commodity hardware, with a ViT-B/16 Vision Transformer serving as the policy’s vision backbone, and reconstruct a Franka Emika manipulator equipped with a custom electric gripper inside NVIDIA Isaac Sim, using the Lula inverse kinematics solver to perform kinematic control from the policy generated end effector commands. We identify and formalize the coordinate and action frame transformations required to transfer a policy trained on handheld demonstrations onto a simulated embodiment, and show that the simulated controller tracks the commanded trajectories to subcentimeter accuracy. We further report pick and place rollout statistics across randomized object configurations and identify the visual domain gap between rendered and real observations as the dominant remaining barrier to transfer. These results helped us understand the UMI framework and established a solid foundation for us to move forward toward free floating microgravity manipulation and autonomous dual arm object handover.
Neal D'Andrea, Joshua Wachs, Abdou Wade et al.· National Aerospace and Elect...· 0 citations
Dual-arm manipulation or physical human-robot coordination requires robots to adapt rapidly to changing environments and constraints. Traditional Learning from Demonstration approaches struggle to generalize when faced with out-of-distribution scenarios, requiring costly retraining. We propose a Movement Primitive learning algorithm based on Gaussian Processes, combined with real-time zero-shot adaptation through Pathwise Conditioning. The method encapsulates the predictive uncertainty of the demonstrated movement using heteroscedastic GPs and utilizes an update via Matheron's rule to instantaneously adjust the trajectory to new via-points, without the need to retrain the underlying model. This formulation is extended to dual-arm coordination by dynamically calculating 6D relative constraints to maintain a closed kinematic chain. Experimental results, both in 2D comparisons against task-parameterized models and in tasks with the ADAM robot, demonstrate robust adaptation with near-zero error in real time, making it applicable for highly changing environments.
Adrián Prados, L. Lishan, Alberto Mendez et al.· Jornadas de Automática· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.