G-VTM: A Multimodal Vision-Trajectory Model for Generalized Vehicle Trajectory Prediction
G-VTM, a generalized vision-trajectory model, is proposed, which captures global map semantics while modeling scenario-and direction-aware interaction based on intuitive visual perception and achieves strong generalized performance under heterogeneous traffic conditions.