G-VTM, a generalized vision-trajectory model, is proposed, which captures global map semantics while modeling scenario-and direction-aware interaction based on intuitive visual perception and achieves strong generalized performance under heterogeneous traffic conditions.
A Stochastic Gating Decoder for multimodal latent variable sampling, adaptively fusing kinematics and data-driven paths to capture driver intention uncertainty while maintaining kinematic consistency is introduced.
Highway ramp merging zones represent a typical dense traffic scenario where complex vehicle interactions make accurate trajectory prediction particularly challenging. This paper proposes a novel dual-encoder model with an interactive time horizon mechanism to address these challenges. The model employs independent encoding channels to separately capture individual motion features and multi-vehicle interaction features, while the introduced selector enables dynamic selection of potentially interacting vehicles. Evaluated on the Chinese highway dataset Expressway-SQM2, our approach demonstrates significant improvements: the dual-encoder architecture alone reduces 5-second prediction RMSE by 29.4% compared with the baseline model, while the complete model with this mechanism achieves a 34.5% reduction. Experimental results confirm the model's effectiveness in handling complex merging scenarios and its strong generalization capability across varying traffic conditions.
Shuai-Peng Liu, Guodong Yin, Weichao Zhuang et al.· 2026 3rd International Confe...· 0 citations
End-to-end autonomous driving in urban environments requires robust decision-making under partial observability and complex multi-agent interactions. Severe occlusions and dense traffic at intersections limit the perception capability of single-agent systems, motivating recent efforts on Vehicle-to-Infrastructure (V2I) cooperation for perception and planning. However, existing evaluation protocols face a fundamental trade-off: open-loop evaluation fails to capture error accumulation and recovery from deviations, while closed-loop evaluation is costly, difficult to scale, and often relies on simulated environments that may suffer from domain gaps. To bridge this gap, we propose VIPS, a benchmark for cooperative autonomous driving in V2I settings based on pseudo-simulation. VIPS extends pseudo-simulation by integrating vehicle and infrastructure observations. This enables scalable yet realistic evaluation of robustness and error propagation without full simulation. We further present CoS-V2X, a cooperative planning framework based on sparse representations. CoS-V2X models vehicle-infrastructure interactions using compact features for efficient communication and robust decision-making under heterogeneous observations. Code and dataset are available at https://vips2026.github.io.
Hoonhee Cho, Jae-Young Kang, Giwon Lee et al.· 0 citations
Road-constrained trajectory recovery from traffic surveillance videos has become a critical component in applications ranging from intelligent transportation to urban planning. Existing approaches typically perform explicit camera-road matching, which either relies on labor-intensive camera calibration or suffers from matching errors. Moreover, prior methods often represent trajectories as sequences of tuples (e.g., node or edge with timestamps), which limits effective spatio-temporal mining, leading to suboptimal recovery accuracy. To address these problems, we propose a novel trajectory recovery framework based on spatio-temporal voxel representation. The real-world coordinates of trajectory nodes are encoded as voxel indices, which naturally capture spatio-temporal relationships. Our framework leverages this voxel representation to enable direct road-constrained trajectory recovery from camera observations, eliminating the need for explicit heuristic camera-road matching. Building upon this voxel formulation, we further present a shape-aware evaluation scheme that assesses trajectories within a spatio-temporal voxel grid, capturing edge and structural information to provide a complementary perspective to existing evaluation metrics. Finally, extensive experiments across multiple datasets demonstrate that our method achieves state-of-the-art trajectory recovery accuracy, with an average F1-score improvement of 14.7%.
Taihang Dong, Jun Zhang, Ping Chen et al.· Proceedings of the 32nd ACM...· 0 citations
Accurate trajectory prediction is critical for autonomous driving safety and energy-efficient motion planning in sustainable urban mobility. This study isolates the effect of local lane graph conditioning by proposing a waterflow method that extracts an ego-centric lane topology via breadth-first traversal of the HD map (For convenience, all acronyms used throughout this manuscript (e.g., HD, LSTM, ADE, FDE, minADE, minFDE, BFS, GNN, SOTA) are collected in the Abbreviations section at the end of the paper), fusing lane features into trajectory encoders through cross-attention. The evaluation is deliberately scoped to ego-vehicle prediction at signal-controlled intersections: we evaluate across two architectures (LSTM and Transformer), two horizons (3 s and 8 s), and both single- and multi-modal (K=6) settings on 89,258 such scenarios from the Waymo Open Motion Dataset, so that “generality” refers to consistency across architectures, horizons, and output settings within this scope, rather than across prediction tasks. Lane conditioning consistently improves accuracy: +9.3% ADE at 3 s (p=0.007, 3 seeds), +26.6% minADE at 8 s (K=6, p=0.003, 3 seeds), and +26.8% ADE for the Transformer (p=0.030, 3 seeds)—with only ∼8% additional parameters for the LSTM. A controlled full-graph ablation (nearest 64 lanes) shows a consistent but not statistically significant trend favouring topologically guided local selection over brute-force spatial proximity (+11.4% minADE, p=0.063). Error decomposition reveals balanced lateral (+26.5%) and longitudinal (+25.4%) improvements, and a per-maneuver analysis over all 13,388 validation scenarios shows the largest gains for turning maneuvers. The lane-conditioned model (<700,000 parameters) runs in 0.8 ms per prediction on a desktop GPU (0.7 ms single-threaded CPU) with below 30 MB peak memory and an estimated 51 mJ per prediction, suggesting feasibility for resource-constrained deployment, pending validation on production automotive hardware.
Xing-Nan Zhou, C. Alecsandru· Sustainability· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.