Skip to content
Preprint

PedestrianDiffusion: Multimodal Generative Denoising and Dense State Estimation for Inertial Navigation

Jul 2026 · 0 citations · 35 references
Computer Science

TL;DR

This work proposes PedestrianDiffusion, a multimodal spectral-domain generative framework reformulating dense 6D state estimation as a continuous conditional denoising process, and introduces a zero-shot semantic conditioning mechanism leveraging vision-language embeddings as categorical priors to generalize across heterogeneous sensor noise profiles.

Abstract

The accuracy of consumer-grade inertial navigation is bottlenecked by the stochastic noise of Micro-Electro-Mechanical Systems (MEMS). Traditional deterministic neural architectures often succumb to ``estimation jittering,''sacrificing high-frequency kinematic fidelity for numerical stability. We propose PedestrianDiffusion, a multimodal spectral-domain generative framework reformulating dense 6D state estimation as a continuous conditional denoising process. By operating in the frequency domain, our formulation bounds the spectral covariance, acting as a mathematical preconditioner to stabilize the reverse diffusion trajectory. Furthermore, we introduce a zero-shot semantic conditioning mechanism leveraging vision-language embeddings as categorical priors to generalize across heterogeneous sensor noise profiles. To address the computational intractability of generative tracking, we deploy a single-step deterministic probability flow ODE solver ($T=1$). This yields high-capacity asynchronous batch trajectory refinement, establishing the viability of generative architectures for asynchronous batch trajectory refinement on edge hardware. Extensive evaluations on the OxIOD, RIDI, RoNIN, and TLIO benchmarks demonstrate that PedestrianDiffusion achieves state-of-the-art performance, exhibiting unprecedented robustness to impulse perturbations and coupled 6D kinematic drift. This work provides a rigorous algorithmic blueprint for next-generation Neural Inertial Measurement Units (N-IMUs).

View source

Similar papers

Book Aug 2026

RAPID: A Scalable and Controllable Physics-Informed Diffusion Framework for Real-Time Pedestrian Trajectory Generation

Generating realistic and diverse pedestrian background flows is critical for numerous downstream applications, ranging from the training and validation of autonomous driving systems to the simulation of mobile communication networks. While recent diffusion-based models achieve state-of-the-art accuracy, they suffer from prohibitive inference latency, lack of physical consistency, and an inability to generalize across heterogeneous datasets, rendering them impractical for industrial Hardware-in-the-Loop testing. To address these challenges, we propose Real-time Adaptive Physics-Informed Diffusion (RAPID), a unified framework explicitly designed to balance high-fidelity generation with strict real-time constraints. First, we introduce a Canonical Representation Module that harmonizes diverse datasets via coordinate-invariant encoding and adaptive modality imputation, enabling unified training across varying scene scales. Second, we propose a Map Context Encoder that decouples computationally expensive map perception from the iterative denoising loop using cached latent embeddings. Third, a Physics-Informed Implicit Sampler integrates Social Force Model gradients as directional priors, encouraging physical consistency (e.g., collision avoidance). Extensive experiments on five heterogeneous benchmarks demonstrate that RAPID establishes a new state-of-the-art balance between fidelity and safety. Notably, it is the only framework capable of operating consistently below the 30 ms industrial threshold, maintaining almost constant inference latency regardless of crowd density. The system exhibits precise controllability over agent behaviors and has been successfully deployed in a production-grade autonomous-driving simulation platform as a core digital twin kernel for large-scale autonomous driving validation. The code is publicly available at https://github.com/tsinghua-fib-lab/RAPID.

Zihan Yu, Huandong Wang, Jingtao Ding et al. · 0 citations
Conference Open access 2026

Xvins: Boosting State Estimation Robustness via Hybrid Temporal Tracking and Efficient Deep Feature Extraction

This work proposes XVINS, a hybrid VIO frontend integrating XFeat—a lightweight deep feature extractor—into the optimization-based VINS-Fusion framework, presenting XVINS as a viable, real-time state estimation solution for agile Micro-Aerial Vehicles (MAVs) and mobile platforms.

Thura Peou, Sarot Srang, Lychek Keo · 0 citations
Sep 2026

AdaStart: Uncertainty-Aware Adaptive Warm Starting for Diffusion-Based Robotic Control

Diffusion policies have demonstrated excellent performance in robotic control tasks, yet their reliance on 50 to 100 denoising steps impedes real-time deployment. Existing acceleration methods based on sampler improvements lack responsiveness to perceptual quality and cannot adaptively adjust computation. Moreover, standard diffusion policies often overlook the reliability of visual observations, leading to degraded robustness under occlusion or domain shift. To address this, we propose AdaStart, an uncertainty-aware adaptive warm starting method, employing a lightweight visual multitask auxiliary network to predict robot state and inform action generation. By estimating visual uncertainty from state prediction error, AdaStart dynamically selects the starting timestep of the reverse diffusion process and constructs a warm-start noisy action via the analytical forward diffusion of the predicted action. With minimal parameter overhead, it accelerates inference under high-confidence conditions while safely reverting to the full diffusion trajectory when uncertain. Experiments on multiple simulated and real-world manipulation tasks show that AdaStart reduces inference steps by nearly 40% while maintaining or improving task success and robustness.

Qi Chen, Xinyang Ren, Jiajun Xing et al. · 0 citations
Conference Aug 2026

Asynchronous Visual-Inertial Motion Parsing via Sparse Differential Dynamics for Agile Mobile Platform

Accurate motion perception is crucial for agile mobile platforms navigating in dynamic environments. However, visual-inertial systems frequently suffer from dominant ego-motion interference and the inherent asynchronous sampling rates between visual sensors and inertial measurement units (IMUs). In this paper, we propose a novel framework for asynchronous visual-inertial motion parsing based on sparse differential dynamics. To alleviate computational bottlenecks, we first employ a sparse visual activation strategy that efficiently extracts high-value features and suppresses redundant background information. To address the irregular and asynchronous sensor sampling, we formulate the evolution of the latent motion field as a continuous-time Neural Controlled Differential Equation (NCDE). This continuous representation acts as a robust temporal prior, decoupling global ego-motion from independent object dynamics. Real-world unmanned aerial vehicle (UAV) flight experiments demonstrate that the proposed framework effectively handles frame loss and aggressive maneuvers. By strictly enforcing physical consistency, our method significantly reduces trajectory drift, offering a lightweight and robust solution for agile mobile platform.

Xin Zhao, Peng Peng, Yonghao Lai et al. · 0 citations
Conference Aug 2026

Robust and efficient UAV tracking via adaptive depth gating and prompt-guided autoregressive decoding

DAR-Track is proposed, a novel framework that harmonizes dynamic computation with generative modeling and outperforms state-of-the-art methods, including MixFormer and SGLATrack, while maintaining superior inference speeds suitable for real-time aerial robotics.

Wenqin Dong · 0 citations
Preprint Jul 2026

BridgeFlow: Fast and Robust SE(2)-Equivariant Motion Planning with Flow Matching

In robotic motion planning, equivariance to rigid body transformations is crucial for robust spatial generalization. However, current learning-based planners face a critical dilemma: they either lack inherent equivariance, treating transformed tasks as novel scenarios, or enforce it via computationally expensive specialized architectures that bottleneck real-time inference. To break this trade-off, we propose BridgeFlow, a fast and strictly SE(2)-equivariant generative motion planning framework. Rather than relying on heavy equivariant networks, BridgeFlow achieves exact spatial equivariance via a lightweight task-centric canonicalization module, enabling generalization using standard architectures. To further accelerate inference, we pair a Brownian bridge informative prior with context-aware mini-batch optimal transport. This constructs a straightened vector field that minimizes transport costs and stabilizes training. Furthermore, environmental awareness is explicitly embedded via Classifier-Free Guidance. Evaluations in dense 2D environments and on a 7-DoF Franka manipulator demonstrate that BridgeFlow achieves up to a 15x inference speedup and a 2x higher valid trajectory rate over state-of-the-art diffusion baselines, alongside robust generalization to entirely unseen environments and arbitrary spatial transformations.

Xinzhe Zhou, Xuyang Wang, Xiaoming Duan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.