Skip to content
Open access

Embodied cognition-driven interpretable trajectory prediction of autonomous systems

Jul 2026 · Nature Communications · Vol 17 · 0 citations · 85 references
Medicine

TL;DR

The study presents an interpretable method for predicting future paths of vehicles and pedestrians by combining scene-level attention, graph-based social interaction reasoning, and physics-based motion constraints, improving accuracy, safety, and speed while enabling transparent decisions.

Abstract

For autonomous systems to operate safely and reliably in dense traffic, they must perform trajectory prediction with human-like, interpretable reasoning. Prevailing data-driven “black-box” models fundamentally lack this capability. This research proposes a paradigm shift toward embodied intelligence, unifying cognitive science principles into a hierarchical framework: a Scene Attention Mechanism for threat prioritization, Social Impact Theory-driven graphs for intent inference, and a physics-compliant Social Force Model. Experimental results demonstrate that our framework reduces average displacement error by 42% and Final Displacement Error by 40% compared to existing state-of-the-art models on ETH and UCY, while enabling near-real-time inference (0.003 s). Crucially, the model’s interpretable architecture, which is validated through risk-sensitive heatmaps and graph visualizations, reveals how agents dynamically balance safety, efficiency, and socio-cultural norms. Beyond performance gains, this work constructs an interpretable bridge between computational models and human cognitive science, laying a foundation for trustworthy autonomous systems. The study presents an interpretable method for predicting future paths of vehicles and pedestrians by combining scene-level attention, graph-based social interaction reasoning, and physics-based motion constraints, improving accuracy, safety, and speed while enabling transparent decisions.

Read PDF

Similar papers

Preprint Aug 2026

Toward the Cognitive--Physical Limits of Embodied Intelligence through a World-Model-Centric Autonomous Racing Agent

Embodied artificial intelligence aims to develop agents that perceive, reason, and act through continuous interaction with the physical world. However, most embodied systems are still evaluated within conservative safety margins or moderate interaction regimes, leaving their capability boundaries under extreme conditions insufficiently understood. Autonomous racing provides a stringent testbed by combining high-frequency localization and perception, adversarial interaction, near-saturated vehicle dynamics, and strict safety constraints. Existing systems push high-speed performance but rarely model and refine cognitive and physical limits jointly. Here we show that a world-model-centric autonomous racing agent provides a concrete step toward exploring these coupled limits. The framework learns predictive world models from near-limit successes and failures to capture interaction evolution, ego dynamics, and feasible-motion boundaries, coupling world-state construction, future-aware reasoning, and near-limit control in a closed-loop refinement process. Training data were collected from real-vehicle autonomous racing, where the onboard system maintained robust localization and perception at speeds up to 256.3 km/h and peak lateral acceleration of 26.8 m/s$^2$. In full-scale simulated racing, the well trained world-model-centric agent achieves an 88.3% interaction success rate across various challenging simulated racing scenarios. Closed-loop refinement of the world model and policy further improved utilization of cognitive-physical limits, recovery from failure modes, and generalization across varying conditions and unseen circuits. These results suggest a boundary-aware methodology in which world models help embodied agents represent, predict, and continually refine their capability boundaries for safer real-world deployment.

Zitong Shan, Baichuan Lou, Yan-Xin Zhou et al. · 0 citations
Open access Jul 2026

An Interpretable and Edge Deployable Spatio-Temporal Trajectory Prediction for Autonomous Driving

A comprehensive Explainable AI (XAI) evaluation framework is introduced, including temporal sensitivity analysis, interaction-aware perturbation studies, spatial influence analysis, and gradient-based feature attribution methods that provide insights into how the model captures temporal motion dependencies, neighboring vehicle interactions, and environmental context during trajectory prediction.

R. Megalingam, Naveen Prasaad Selvarajan, Pritty Vijay · 0 citations
#reinforcement learning Review Open access Sep 2026

A survey of world models for physical AI with uncertainty representation and control

Physical AI systems must reason about real-world dynamics in order to perceive, predict, and act safely under partial observability and uncertainty. World models–learned predictive representations of environment dynamics and action consequences–have emerged as a unifying framework for integrating perception, prediction, planning, and control in embodied agents. This survey provides a comprehensive and technically grounded review of learning-based world models for Physical AI, with particular emphasis on closed-loop decision-making. We organize existing approaches along six compositional design dimensions: state abstraction, temporal dynamics, uncertainty source and treatment, structural prior, observation modality, and decision coupling. Beyond this design-oriented taxonomy, we analyze how world models interact with optimization–highlighting compounding error, planner exploitation, rollout horizon management, and uncertainty calibration as central design tensions. We further examine evaluation methodologies, benchmark ecosystems, and sim-to-real transfer challenges, and synthesize open problems in long-horizon consistency, physical constraint enforcement, data efficiency, and safety. By clarifying recurring trade-offs across robotics and model-based reinforcement learning, this survey outlines principled directions for building reliable and scalable Physical AI systems.

Sven Kirchner, Nils Purschke, Alois Knoll · 0 citations
Preprint Aug 2026

HarnessWAM: Bridging Prediction and Deliberation in World Action Models

Results demonstrate that model-external structured state maintenance and closed-loop agentic decision making can effectively extend the local control capabilities of WAMs into embodied task execution that is plannable, verifiable, and recoverable.

Zhaopeng Gu, Bingke Zhu, Tianxin Lin et al. · 0 citations
Preprint Aug 2026

Learning the Right Abstraction: Neural Reduced Dynamics for Complex Robot Control

A neural reduced dynamics framework is developed that separates the state the model propagates from what can be supplied as an input or recovered analytically, trains policies entirely inside the frozen learned model, and validates them back in the high-fidelity simulator.

Harry Zhang, Dan Negrut · 0 citations
Open access Aug 2026

ESIM: An Embodied System Integration Methodology for Real-Time Risk Mitigation in Autonomous Driving

Traditional modular pipelines in autonomous driving (AD) frequently suffer from error accumulation and delayed responsiveness during safety-critical events. Although Embodied Intelligence (EI) introduces a paradigm shift through internal “World Models” for proactive risk mitigation, a substantial gap remains between high-level cognitive theories and real-time, safety-certified deployment. This paper bridges that gap by proposing an Embodied System Integration Methodology (ESIM), which translates cognitive models into fielded robotic systems. Grounded in a “Perception-Imagination-Execution” (PIE) cognitive architecture, ESIM treats risk prediction as an uncertainty-driven, counterfactual closed-loop sensorimotor process. Unlike passive prediction models, the framework employs a Bayesian uncertainty-gated mechanism that selectively triggers a World Model to simulate future risk scenarios only when perceptual degradation occurs. We validate this methodology through a multi-paradigm study spanning three distinct levels: an academic prototype on edge computing platforms, an industrial implementation adhering to ASIL-D (Automotive Safety Integrity Level D) constraints, and an open-source simulation platform. The results demonstrate that by applying hardware acceleration and asynchronous pipelines, the ESIM framework consistently maintains end-to-end latencies within 10–20 ms across heterogeneous hardware. We explicitly address the engineering trade-offs in latency, hardware heterogeneity, and optimization, and establish mathematically grounded probabilistic safety boundaries for black-box neural architectures. Finally, we discuss the framework’s scalability in extreme scenarios, coupling with SLAM pipelines, privacy-preserving federated learning, and generalization potential in the low-altitude economy.

Daiquan Xiao, Qi-Hao Liu, Xuecai Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.