Skip to content

MaTF: Maneuver-Aware Temporal Fusion for Trajectory Prediction Under Arbitrary Observation Length

Oct 2026 · IEEE Robotics and Automation Letters · Vol 11, pp. 11078-11085 · 0 citations · 35 references

Abstract

Trajectory prediction is essential for many robotic applications, yet most existing models rely on fixed-length observations and struggle with temporally irregular inputs. In real-world settings, prediction difficulty further increases when agents exhibit strong maneuverability, as their future motions depend on distinct short-term and long-term temporal cues. A Maneuver-aware Temporal Fusion framework is proposed to separate short-term dynamics from long-term intentions and fuse them through a motion-complexity-guided attention mechanism. The framework first extracts temporal features at different scales, and then adaptively balances them according to the maneuver patterns of each agent. To support incomplete or short observations, a self-distillation strategy is introduced to reconstruct missing motion segments, enabling consistent prediction without relying on explicit teacher-student models. Furthermore, a Mamba-Transformer hybrid backbone is employed to enhance computational efficiency and improve generalization under arbitrary observation lengths. Experiments on the ETH/UCY and SDD datasets show that MaTF consistently outperforms existing methods, particularly in scenarios with irregular or shortened observations.

View source

Similar papers

W2W: Language-Model-Based Trajectory Prediction with Reinforcement Learning

This work converts observed trajectories and interaction cues into parsable textual prompts, so that interaction semantics are expressed more explicitly in the model input and remains competitive with recent LM-based prediction methods and strong trajectory prediction base-lines on ADE/FDE.

Zi-Rui Xu, Biao Yang, Rongrong Ni et al. · 1 citation
Jul 2026

Context-Informed Ship Trajectory Prediction via Conditional Attention

Long-term ship trajectory prediction is a fundamental capability for maritime safety and autonomous navigation. While recent Transformer-based architectures have improved forecasting horizons, they predominantly rely on historical kinematic states, treating vessel motion as an isolated system. In reality, maritime navigation is profoundly modulated by extrinsic factors like weather and constrained by static vessel characteristics. Existing multimodal approaches fundamentally model the joint distribution over states and contexts, treating environmental variables as peer features rather than encoding the directional physical dependence of vessel dynamics on environmental conditions. In this work, we propose the Conditional Informer, a novel encoder-decoder architecture that formulates trajectory prediction as a conditional generation task. We employ a dedicated Conditional Attention mechanism where the vessel state explicitly queries environmental contexts through cross-attention, encoding the physical prior that weather modulates - but is not generated by - vessel dynamics. Furthermore, to address the intermittency of real-world data, we introduce a Modality Masking training strategy to prevent catastrophic degradation during sensor fallback. Extensive experiments on AIS and ERA5 data demonstrate that our approach outperforms kinematic and concatenation-based baselines by 15.4% in prediction accuracy when context is available. Crucially, Modality Masking prevents shortcut learning, reducing fallback error by nearly an order of magnitude compared to unconstrained models.

Yuansheng Guan, C. Squires, Timothy Hu et al. · 0 citations
Preprint Aug 2026

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

SimWAM, a simple yet effective WAM that leverages future-video prediction as a training-time supervision signal, co-trains a pretrained video expert and a lightweight action expert with joint flow matching and applies reinforcement learning to optimize a compositional driving reward beyond trajectory imitation.

Zongchuang Zhao, Xin Zhou, Tianyang Xu et al. · 1 citation · ⚡1
Oct 2026

Asynchrony-Robust Cooperative Perception and Prediction via Continuous-Time Global State Evolution

Vehicle-to-everything (V2X) collaboration can alleviate the limited perception range and occlusion issues of single-agent autonomous driving. However, most existing cooperative studies still focus on single-frame perception, while the few works on joint cooperative perception and prediction largely rely on fixed-step, frame-aligned fusion, making them difficult to apply in realistic systems with cross-agent temporal misalignment. To address this issue, we propose CoSPACE, a novel framework that formulates asynchronous cooperative perception and prediction as continuous evolution and event-triggered correction of a shared scene state. Specifically, CoSPACE maintains an ego-centric global BEV state and propagates it to arbitrary observation timestamps using an ODE-style dynamics model, thereby preserving temporal coherence under irregular multi-agent inputs. When asynchronous observations arrive, they are treated as local evidence and assimilated by an event-triggered update module, which first selects informative regions and then performs time-aware gated correction to adaptively incorporate temporally reliable evidence. With this design, CoSPACE moves beyond the conventional frame-aligned fusion paradigm and mitigates the spatial misalignment, semantic confusion, and prediction degradation caused by temporal asynchrony. Experiments on V2XPnP-Seq under diverse delay settings show that CoSPACE consistently outperforms representative baselines and achieves strong robustness in asynchronous cooperative perception and prediction.

Han-Xiao Ren, Keqiang Li, Xiang Zhao et al. · 0 citations
2026

Hybrid State Space Modeling for Sequence-Based Robot Localization Under Challenging Environments

Visual localization is vital for autonomous systems but remains challenging under dynamic conditions. Transformers offer strong temporal modeling at quadratic cost, while CNNs are efficient yet limited in long-range dependencies. Existing methods also lack robustness to illumination, weather, and seasonal changes, constraining real-world applicability. To address this, this paper proposes AdapseqNet, a dual-branch architecture that integrates stabilized state-space modeling with differential temporal enhancement. First, a stabilized state-space formulation featuring Lyapunov-constrained parameterization and adaptive discretization is proposed, ensuring asymptotic stability and linear computational complexity for reliable processing of extended sequences. Second, a selective Mamba architecture is developed to combine temporal-state modeling with content-aware gating, enabling adaptive feature selection that emphasizes discriminative cues while suppressing redundancy. Third, a differential enhancement module is designed to extract motion-invariant representations through symmetric temporal differencing and LSTM-based refinement, enhancing resilience to appearance variations caused by lighting, weather, and seasonal changes. Beyond architectural design, multi-scale feature fusion and output distribution control are incorporated to optimize representation quality and ensure consistency for similarity-based retrieval. Extensive experiments on multiple benchmarks demonstrate that AdapseqNet achieves a better localization accuracy across diverse and challenging conditions. Note to Practitioners—Visual localization is crucial for autonomous robots but often fails under varying lighting, weather, or seasonal conditions. We propose a dual-path approach: one path captures long-term patterns using control-inspired stable modeling, while the other extracts motion cues that remain consistent despite appearance changes. This combination enables accurate place recognition even in extreme environments. Our system operates efficiently on standard hardware and was tested on an indoor robot, achieving centimeter-level accuracy. This approach can enhance existing navigation systems without requiring additional sensors. Future work will focus on real-time optimization for outdoor deployment.

Zhenyu Li, Tian-Yi Shang · 0 citations
Preprint Jul 2026

SegDiff: Segmented Trajectory Diffusion for Consistent and Adaptive Robot Manipulation

Imitation learning enables robots to acquire manipulation skills from demonstrations by mapping observations to actions. Existing approaches predict either short-horizon continuous action sequences or discrete keyposes. However, continuous prediction methods suffer from compounding errors due to short prediction horizons and struggle with multi-modal action distributions, whereas keypose-based methods necessitate an external planner, constraining real-time applicability. To address these challenges, we introduce SegDiff, a closed-loop visuomotor policy that integrates the strengths of both paradigms. SegDiff decomposes demonstrations into motion segments between keyposes and learns to predict the continuous trajectory from the current state to the next keypose, enabling long-horizon prediction with real-time refinement. Furthermore, we leverage the capability of diffusion models and DDIM inversion to propose a Dynamic Temporal Ensembling mechanism, which allows the policy to efficiently respond to dynamic environments and mitigate discontinuities caused by inconsistent multi-modal sampling. SegDiff demonstrates significant performance gains over existing approaches across various simulated and real-world scenarios, indicating its strong ability to reason over extended temporal dependencies while maintaining real-time adaptability and control stability.

Haidong Cao, Wenjun Cao, Quanhao Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.