2026· E3S Web of Conferences· 0 citations· 3 references
TL;DR
This work proposes XVINS, a hybrid VIO frontend integrating XFeat—a lightweight deep feature extractor—into the optimization-based VINS-Fusion framework, presenting XVINS as a viable, real-time state estimation solution for agile Micro-Aerial Vehicles (MAVs) and mobile platforms.
Abstract
Robust state estimation in GPS-denied environments remains a primary challenge for Visual-Inertial Odometry (VIO). Traditional VIO pipelines relying on hand-crafted features often experience drift or tracking failure under rapid motion and dynamic lighting. While deep learning-based local features offer improved robustness, high computational costs frequently limit their adoption in real-time, resource-constrained systems. In this work, we propose XVINS, a hybrid VIO frontend integrating XFeat—a lightweight deep feature extractor—into the optimization-based VINS-Fusion framework. The system employs a dual-strategy tracking mechanism: computationally efficient KLT optical flow is utilized for high-frequency temporal tracking, while XFeat in-ference is dynamically triggered for feature replenishment during tracking degradation. We evaluate this architecture across the 11 sequences of the EuRoC MAV dataset. The results indicate improved robustness to motion blur, reducing absolute trajectory error by up to 82.7% in highly dynamic scenarios compared to baseline methods. Furthermore, we analyze the theoretical and practical tradeoffs between deep feature quantization and classical sub-pixel precision, presenting XVINS as a viable, real-time state estimation solution for agile Micro-Aerial Vehicles (MAVs) and mobile platforms.
KLTNet is proposed, a lightweight learning-based, plug-and-play sparse feature tracker designed to replace classical KLT trackers in KLT-based VIO front ends and predicts anisotropic confidence weights supervised through differentiable multi-view triangulation, which can be used as observation weights in compatible VIO estimators.
DAR-Track is proposed, a novel framework that harmonizes dynamic computation with generative modeling and outperforms state-of-the-art methods, including MixFormer and SGLATrack, while maintaining superior inference speeds suitable for real-time aerial robotics.
Wenqin Dong· International Conference on...· 0 citations
Monocular visual odometry is important in autonomous driving, robotics, and related fields, and has attracted increasing attention in computer vision.Traditional geometric methods and end-to-end deep learning methods have achieved promising results in monocular visual odometry, but they still face limitations in temporal consistency, uncertainty representation, and abnormal observation handling.To address these issues, this paper proposes a monocular visual odometry method that combines deep temporal features with a Kalman filtering module based on a state-space model. By explicitly modeling the temporal evolution of motion states in the state space, the method introduces continuity constraints into the pose estimation process. At the same time, the state transition matrix, as well as the process noise covariance matrix and the measurement noise covariance matrix, are learned by neural networks, which enables the system to adaptively adjust the fusion weight between prediction and observation. Experimental results on the KITTI VO dataset show that, compared with the baseline method, the proposed model improves pose estimation robustness and adaptability to incomplete observations, which verifies the effectiveness of combining classical filtering theory with deep learning for monocular visual odometry.
Mingyue Wang· Poster Volume 0007 The 2026...· 0 citations
Traditional VIO suffers from optical flow tracking failures and severe localization drift at low frequencies, primarily due to large disparities between consecutive frames. To resolve this issue, a tightly-coupled stereo VIO system driven by deep features was constructed. Instead of a conventional optical flow front-end, the SuperPoint network is utilized to extract semantic features, paired with a brute-force descriptor matching strategy for large-parallax scenarios. Systematic down-sampling simulations (from 20Hz to 1Hz) were performed on the EuRoC dataset. The quantitative results indicate that even at an extreme 1Hz frequency, the proposed system maintains a complete trajectory estimate, demonstrating remarkable survivability. Furthermore, a non-monotonic performance trend was observed. At a 2Hz sampling rate, the strong geometric constraints provided by the wide baseline effectively mitigated IMU integration drift. Consequently, the localization accuracy (RMSE of 0.21m) outperformed the 5Hz midfrequency range and even surpassed the 20Hz baseline. This finding challenges the conventional assumption that higher frequencies inherently yield better accuracy, offering a fresh theoretical perspective for designing robot navigation systems in resource-constrained environments.
Ruizhe Chen· International Conference on...· 0 citations
Visual localization is vital for autonomous systems but remains challenging under dynamic conditions. Transformers offer strong temporal modeling at quadratic cost, while CNNs are efficient yet limited in long-range dependencies. Existing methods also lack robustness to illumination, weather, and seasonal changes, constraining real-world applicability. To address this, this paper proposes AdapseqNet, a dual-branch architecture that integrates stabilized state-space modeling with differential temporal enhancement. First, a stabilized state-space formulation featuring Lyapunov-constrained parameterization and adaptive discretization is proposed, ensuring asymptotic stability and linear computational complexity for reliable processing of extended sequences. Second, a selective Mamba architecture is developed to combine temporal-state modeling with content-aware gating, enabling adaptive feature selection that emphasizes discriminative cues while suppressing redundancy. Third, a differential enhancement module is designed to extract motion-invariant representations through symmetric temporal differencing and LSTM-based refinement, enhancing resilience to appearance variations caused by lighting, weather, and seasonal changes. Beyond architectural design, multi-scale feature fusion and output distribution control are incorporated to optimize representation quality and ensure consistency for similarity-based retrieval. Extensive experiments on multiple benchmarks demonstrate that AdapseqNet achieves a better localization accuracy across diverse and challenging conditions. Note to Practitioners—Visual localization is crucial for autonomous robots but often fails under varying lighting, weather, or seasonal conditions. We propose a dual-path approach: one path captures long-term patterns using control-inspired stable modeling, while the other extracts motion cues that remain consistent despite appearance changes. This combination enables accurate place recognition even in extreme environments. Our system operates efficiently on standard hardware and was tested on an indoor robot, achieving centimeter-level accuracy. This approach can enhance existing navigation systems without requiring additional sensors. Future work will focus on real-time optimization for outdoor deployment.
Zhenyu Li, Tian-Yi Shang· IEEE Transactions on Automat...· 0 citations
Abstract. Existing single-object trackers often struggle to fully exploit temporal information and distinct feature representations, limiting their robustness in complex scenarios. Specifically, current approaches face challenges in (1) capturing channel-wise discriminative features during target deformation, (2) adapting to sudden appearance changes due to reliance on static memory banks, and (3) mitigating background interference within transformer attention mechanisms. To address these issues, we propose STFTrack, a proposed tracking framework integrating enhanced spatiotemporal features. First, we construct a semantic–spatial–instance attention module, which refines target representation via cascaded channel calibration and instance uncertainty modeling. Second, a gated attention fusion module is introduced to adaptively aggregate temporal and spatial cues. Third, we design an improved spatiotemporal decoder equipped with learnable autoregressive queries to establish a robust cross-frame state propagation mechanism. Extensive experiments on five benchmarks, including LaSOT, GOT-10k, and TrackingNet, demonstrate the superiority of STFTrack. For instance, it achieves an area under the curve score of 72.8% on the LaSOT benchmark. Qualitative results further confirm that STFTrack effectively suppresses background clutter and maintains accurate tracking under challenging conditions including deformation, fast motion, and occlusion.
Changhao Zhou, Xun Duan, Guangqian Kong et al.· Journal of Electronic Imagin...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.