Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· 0 citations· 15 references
Abstract
Generating realistic and diverse pedestrian background flows is critical for numerous downstream applications, ranging from the training and validation of autonomous driving systems to the simulation of mobile communication networks. While recent diffusion-based models achieve state-of-the-art accuracy, they suffer from prohibitive inference latency, lack of physical consistency, and an inability to generalize across heterogeneous datasets, rendering them impractical for industrial Hardware-in-the-Loop testing. To address these challenges, we propose Real-time Adaptive Physics-Informed Diffusion (RAPID), a unified framework explicitly designed to balance high-fidelity generation with strict real-time constraints. First, we introduce a Canonical Representation Module that harmonizes diverse datasets via coordinate-invariant encoding and adaptive modality imputation, enabling unified training across varying scene scales. Second, we propose a Map Context Encoder that decouples computationally expensive map perception from the iterative denoising loop using cached latent embeddings. Third, a Physics-Informed Implicit Sampler integrates Social Force Model gradients as directional priors, encouraging physical consistency (e.g., collision avoidance). Extensive experiments on five heterogeneous benchmarks demonstrate that RAPID establishes a new state-of-the-art balance between fidelity and safety. Notably, it is the only framework capable of operating consistently below the 30 ms industrial threshold, maintaining almost constant inference latency regardless of crowd density. The system exhibits precise controllability over agent behaviors and has been successfully deployed in a production-grade autonomous-driving simulation platform as a core digital twin kernel for large-scale autonomous driving validation. The code is publicly available at https://github.com/tsinghua-fib-lab/RAPID.
This thesis addresses fundamental challenges in traffic scene prediction for autonomous driving by introducing robust and computationally efficient models based on polynomial representations. While conventional sequence-based representations often struggle with noise and generalization, this work demonstrates that polynomial representations offer significant advantages in computational efficiency, generalization, and prediction plausibility. Through theoretical analysis and empirical validation, this thesis demonstrates that moderate-degree polynomials capture real-world motion dynamics with high fidelity without constraining predictive performance. Building on this foundation, a prediction model representing both trajectories and map geometry with polynomial representations achieves near state-of-the-art accuracy on standard benchmarks while substantially improving generalization under distribution shift. Extending this concept, a diffusion- based generative framework enables multi-agent scene generation, producing traffic continuations that are more plausible and kinematically consistent than those generated by conventional baselines. Evaluations on the Argoverse 2 and Waymo Open datasets confirm that polynomial representations reduce computational cost, enhance cross-dataset generalization, and yield smoother trajectories and higher behavioral plausibility. The findings reveal that standard in-distribution evaluation and regression-based metrics may fail to reflect true model generalization and prediction plausibility. By providing theoretical justification and empirical validation, this dissertation estab- lishes polynomial trajectory representations as an efficient, expressive, and generalizable foundation for traffic scene prediction in safety critical autonomous driving.
GeoFlow is a novel framework designed to achieve efficient driving video generation by harnessing explicit geometric priors, using a Geometry-Aligned Prior (GAP) distribution as starting point, and can achieve remarkable efficiency of both training and inference.
Jiazheng Liu, Hangbiao Li, J. Zhang et al.· 0 citations
Human trajectory prediction has significant practical applications in various scenarios, such as autonomous driving, social robots and so on. Recently, it has been widely studied by diffusion models in order to model the inherent multi-modality of human motions. However, existing diffusion-based approaches only focus on modeling the social interactions
via
a single encoder and neglect the scene interactions, which results in producing unreasonable trajectories across obstacles or road boundaries. To address this issue, we propose the Interaction-Aware Diffusion Model (IADM), a novel diffusion-based framework considering both human motions and surrounding scene layout by treating the social and scene interactions as conditions in the parameterized reverse Markov chain. To implement IADM, we design two encoders,
i.e
., social encoder and scene encoder, where the social encoder models the social interactions
via
attention mechanism, and the scene encoder preserves spatial information of the scene when learning the scene interactions. Furthermore, we devise the dual-guidance decoder consisting of the motion-guided temporal module and the scene-guided spatial module to intensify the collaboratively guidance of the social and scene interactions. Extensive experiments on the ETH/UCY dataset, Stanford Drone Dataset and Intersection Drone Dataset validate the superiority of our method, achieving state-of-the-art results.
Zhong Zhang, Nuoran Wang, Song Gao et al.· PeerJ Computer Science· 0 citations
This work proposes FlashMo, a frequency-aware sparse motion diffusion model that prunes low-frequency tokens to enhance efficiency without custom kernel design, and introduces MotionSiT, a scalable diffusion transformer based on a joint-temporal factorized interpolant with Lie group geodesics over SO(3) manifolds, enabling principled generation of joint rotations.
Zeyu Zhang, Yiran Wang, Danning Li et al.· Neural Information Processin...· 11 citations
Accurate multi-agent trajectory prediction requires modeling inherently multi-modal futures while maintaining physical and social plausibility. Diffusion-based predictors provide strong generative coverage but can still produce unsafe samples such as inter-agent collisions, especially under short observation horizons. We propose an energy-guided diffusion framework that injects interpretable constraints directly into the denoising process at inference time. Our model first encodes agent interactions into condition tokens and then samples K future trajectories with a conditional diffusion generator. During sampling, we apply per-step gradient guidance derived from a differentiable energy decomposition that penalizes social collisions and encourages intent/goal consistency, enabling plug-and-play controllability without retraining. Experiments on ETH/UCY and SDD under the standard 8 → 12 protocol demonstrate consistent improvements in best-of-K accuracy (minADE/minFDE) and a substantial reduction in collision rate. Ablation studies further show that the social energy term drives most safety gains, the intent term improves long-horizon accuracy, and continuous per-step guidance is essential compared to late-step or post-hoc alternatives.
Hongkun Wang, Fei Li· International Conference on...· 0 citations
Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry drift, and inconsistent object motion. Recent latent video diffusion models have achieved impressive visual quality, but directly applying them to structured traffic scenes often leads to unstable geometry and degraded temporal coherence over extended horizons. We present a training-free inference framework that stabilizes reliable static structure in pretrained video predictions through multi-frame temporal context and view-conditioned routing. For front-camera videos, our method refines generated futures with a multi-frame depth-layered renderer that projects static geometry from observed history frames while preserving dynamic regions from the generative base model. For heterogeneous traffic views, a frozen vision-language model infers a coarse camera group from the observed clip and selects a specialized motion-based predictor. The framework requires neither retraining nor fine-tuning of the underlying video model and can be applied directly to pretrained generators. We validate the proposed framework on the AI City Challenge Track 5 benchmark, where our final system achieves competitive performance among the top-ranked teams. These results demonstrate that geometry-aware inference-time refinement and view-conditioned hybrid inference can improve static-geometry stability and low-level structural fidelity without changing the original model architecture.
K. M. Le, H. Pham, Luu Thanh Danh et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.