Skip to content
Preprint

ORBIT++: Benchmarking SfM in the Wild with 360{\deg} Video

Aug 2026 · 0 citations · 36 references
Computer Science

TL;DR

A new benchmark for evaluating camera pose estimation is introduced, called ORBIT, to leverage online panoramic 360{\deg} video as a source of data from which to construct challenging clips, while still enabling robust ground-truth trajectory recovery.

Abstract

Structure-from-Motion (SfM) is a cornerstone of 3D perception, yet current methods often fail when applied to complex videos involving challenging camera motions or dynamic scenes. Compounding the problem, the field lacks reliable ground-truth benchmarks for such difficult scenarios, making it hard to gauge real-world progress or to pinpoint where improvements are most needed. To address this gap, we introduce a new benchmark for evaluating camera pose estimation. Our key insight is to leverage online panoramic 360{\deg} video as a source of data from which to construct challenging clips, while still enabling robust ground-truth trajectory recovery. The panoramic nature of these videos provides richer visual context for tracking camera motion, even when parts of the view are affected by blur, motion, or dynamic objects. After tracking camera motion across full 360{\deg} videos, we crop and reproject selected portions to generate perspective-view clips that serve as our benchmark, called ORBIT. Experiments show that COLMAP, as well as recent optimization-based and feed-forward SfM methods struggle to accurately estimate camera poses on our benchmark. Hence, ORBIT provides a valuable testbed where researchers can meaningfully measure progress on truly challenging, real-world SfM problems.

View source

Similar papers

Jul 2026

VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion

This work introduces a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos.

Zador Pataki, Paul-Edouard Sarlin, Marc Pollefeys · 0 citations
Preprint Sep 2026

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

Multi-camera systems are increasingly practical for robotics, AR/VR, and autonomous driving because complementary views reduce depth ambiguity and preserve visibility under occlusion. Existing point-tracking benchmarks, however, focus on a single video or static multi-camera rigs. None test long-term 3D point tracking across several synchronized views under camera motion. We introduce TAPVid-MV (Tracking Any Point in Video across Multiple Views), the first benchmark for this setting. It contains a curated set of 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks across seven subsets spanning indoor and outdoor domains, from robotics and human activity to driving and synthetic procedural scenes. We obtain these trajectories using dataset-specific auxiliary modalities: sensor depth, LiDAR, SLAM and SfM points, human meshes, posed object meshes, and simulation. Every sequence and trajectory is visually verified by human annotators. Across more than 30 baselines, no method comes close to solving the task. Surprisingly, existing multi-view point trackers do not consistently outperform monocular point trackers. By evaluating reconstruction and point tracking on the same datasets, TAPVid-MV helps distinguish errors in recovered geometry from errors in point correspondence. Through this joint analysis, we identify geometry recovery as a major bottleneck for accurate 3D point tracking. Beyond multi-view 3D point tracking, our released annotations support monocular 2D and 3D point tracking, future-trajectory prediction, and 4D reconstruction.

Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman et al. · 0 citations
Jul 2026

OmniX: Any-view and Any-time 4D Reconstruction via Feed-forward Trajectory Fields

OmniX achieves state-of-the-art performance on dense 3D point trajectory prediction and 3D point tracking, while also demonstrating competitive results on video depth estimation and camera pose estimation.

Yanqin Jiang, Tengfei Wang, Zhengwei Wang et al. · 3 citations · ⚡1
Jul 2026

TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting

This work proposes TARS, a 3D-free video re-shooting paradigm that provides more accurate and temporally consistent camera control than prior methods, and introduces self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data or 3D reconstruction.

Jiwen Liu, Shujuan Li, Xiaohan Li et al. · 0 citations
Preprint Aug 2026

MV2: Multi-View Multi-Vehicle Driving Dataset for Novel View Synthesis

Differentiable rendering has advanced novel view synthesis (NVS), yet applying it to real-world driving remains difficult due to sparse capture viewpoints, dynamic objects, and limited multi-trajectory data. We introduce the Multi-View Multi-Vehicle (MV2) dataset and benchmark for evaluating NVS models under large viewpoint changes in dynamic urban scenes. MV2 features synchronized captures from a car, scooter, and drone, each following distinct yet synchronized trajectories. Training NVS methods on one vehicle's camera stream and testing on another enables evaluation under substantially larger viewpoint variations than existing single-trajectory datasets. All sequences are registered via Structure-from-Motion and camera poses verified using manual pixel-level correspondence annotations, yielding 50 high-quality scenes with 12000 images. Benchmarking recent NVS and camera pose estimation methods shows that NVS performance degrades with increasing viewpoint disparity, and that feed-forward pose estimators notably lag behind optimization-based approaches, highlighting MV2 as a rigorous testbed for NVS in driving. The dataset, benchmark protocol, and project resources are available at https://mv2-dataset.github.io/.

Sanjay Bhargav Dharavath, Hanvitha Saraswathi Mukkamala, F. Khan et al. · 0 citations

FastEventDGS: Deformable Gaussian Splatting for Fast Dynamic Scenes from a Single Event Camera

This work introduces FastEventDGS, a novel Deformable Gaussian Splatting-based framework that leverages a single event camera for high-fidelity 4D reconstruction in dynamic scenes and proposes a local patch event motion loss to constrain object motion, effectively mitigating over-fitting.

Zijia Dai, Nico Messikommer, Rong Zou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.