Jul 2026· Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Talks· pp. 1-3· 0 citations· 2 references
TL;DR
A LiDAR-Constrained Deep Visual Odometry system, a robust tracking architecture designed to solve these "impossible" shots by fusing pre-existing LiDAR geometry with modern Deep Learning, and introduces a "Leapfrogging" architecture that automatically detects and corrects temporal drift by re-anchoring to the geometry from trusted keyframes.
Abstract
Matchmoving is the bedrock of visual effects, yet it remains a fragile bottleneck when footage contains heavy motion blur, low texture, or dynamic occlusion. While physical on-set camera tracking (e.g., encoded cranes) exists, it is often impractical for handheld interior shots and prone to mechanical slippage, leaving post-production software to solve the gap. This talk presents a LiDAR-Constrained Deep Visual Odometry system, a robust tracking architecture designed to solve these "impossible" shots by fusing pre-existing LiDAR geometry with modern Deep Learning. Unlike traditional commercial solvers that hunt for sparse, high-contrast corners, our approach uses Deep Optical Flow (RAFT) to track the entire dense image context, locking the camera directly to the set’s 3D mesh. We introduce a "Leapfrogging" architecture that automatically detects and corrects temporal drift by re-anchoring to the geometry from trusted keyframes. By prioritizing geometric truth over feature quantity, this standalone Python tool reduces days of manual hand-tracking and rotoscoping to minutes of automated computation, achieving high median precision on sequences where standard algorithms fail entirely.
This work introduces a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos.
Zador Pataki, Paul-Edouard Sarlin, Marc Pollefeys· arXiv.org· 1 citation· ⚡1
Multi-Camera People Tracking (MCPT) traditionally relies on precise intrinsic and extrinsic camera calibration to project 2D detections into a unified 3D world coordinate system.However, manual calibration constitutes a major bottleneck in large-scale dataset generation from unconstrained video archives. This work prop...
This work introduces FastEventDGS, a novel Deformable Gaussian Splatting-based framework that leverages a single event camera for high-fidelity 4D reconstruction in dynamic scenes and proposes a local patch event motion loss to constrain object motion, effectively mitigating over-fitting.
Zijia Dai, Nico Messikommer, Rong Zou et al.· 0 citations
KP-SLAM is proposed, which predicts dense optical flow and paired pointmap priors from a shared representation and incorporates them into the same BA backend and introduces a Depth-Scale-Pose-to-Pointmap (DSPP) objective that relates optimized inverse depth, edge-wise relative scale, and camera pose to paired pointmap...
Song Gao, Xinyu Huang, Zheng Huang et al.· Symmetry· 0 citations
The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting; depth is available only for training and validation, so inference is RGB-only. We use two RGB-only routes as a controlled test of one hypothesis: that cross-view geometric con...
Abdullah Naeem, Anav Katwal, Ayon Dey et al.· 0 citations
Multi-object tracking (MOT) has advanced rapidly in urban surveillance and autonomous driving, yet many trackers rely on ReID- and transformer-based appearance encoders and are designed for standard FoV cameras. These assumptions break down for low-cost omnidirectional deployments, where equirectangular projection intr...
Xin Shu, Meegan Gower, Y. Buckley et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.