Skip to content

ObliCity: A Benchmark and Baseline for Roof-to-Ground Projection Displacement Correction

Jul 2026 · arXiv.org · Vol abs/2607.25210 · 0 citations · 24 references
Computer Science

TL;DR

This work reformulates DragOSM into DragRoof, an ODE-based framework inspired by human annotation behavior, and introduces the Oblique City dataset (ObliCity), the first large-scale benchmark that integrates high-resolution UAV imagery and globally distributed satellite data, covering diverse city morphologies and camera perspectives.

Abstract

Oblique-view urban remote sensing imagery inevitably exhibits geometric projection displacements between building roofs and footprints, leading to significant distortions in spatial structure. Existing approaches either ignore these deformations or handle them implicitly within segmentation-based frameworks, where progress is dominated by general segmentation advances rather than improvements in geometric correction. In this work, we explicitly define roof-to-footprint offset vector (RFOV) extraction as an independent learning task that decouples geometric alignment from semantic segmentation. To support this task, we introduce the Oblique City dataset (ObliCity), the first large-scale benchmark that integrates high-resolution UAV imagery and globally distributed satellite data, covering diverse city morphologies and camera perspectives. Methodologically, we reformulate DragOSM into DragRoof, an ODE-based framework inspired by human annotation behavior. By simulating the continuous process of dragging roofs toward their footprints, DragRoof learns deterministic, geometry-consistent offset fields and adaptively determines convergence through an end token. Extensive experiments on ObliCity demonstrate that DragRoof achieves state-of-the-art RFOV extraction performance, requiring fewer inference steps while delivering superior directional and length accuracy. Our dataset and model establish a principled foundation for studying projection displacement correction in oblique remote sensing imagery. The source code and dataset will be avaliable at https://github.com/likaiucas/DragRoof.

View source

Similar papers

2026

GECNet: A Vectorized Building Footprints Extraction Network Based on Geometric Perception and Vertex Guidance

Extracting building footprints from aerial or satellite imagery remains a significant challenge, particularly in maintaining the geometric regularity of man-made structures. While polygon-based methods offer vectorized representations superior to pixel-based approaches, they often struggle with corner ambiguity and fai...

Wen-Jie Zhao, Xue-Jing Xie, Ze Meng et al. · 0 citations
Review Open access Sep 2026

GFE-Net: Geometry-Enhanced Feature Extraction Network for Semantic Segmentation of Large-Scale LiDAR Point Clouds

Accurate semantic segmentation of large-scale outdoor LiDAR point clouds remains a challenging endeavor, primarily due to ambiguous class transitions at object interfaces, non-uniform sampling density across the surveyed area, and shared geometric signatures among distinct object categories. This paper proposes GFE-Net...

Hui Liu, Guang-Ming Zhang, Chuang Chen et al. · 0 citations
Preprint Aug 2026

The Coastline as a Structural Constraint: Harnessing Scene Geometry for Autonomous Surface Vessel Localization

Coastal environments contain rich, largely unexploited geometric structure capable of providing globally referenced localization cues. In this work, we present two complementary localization frameworks that exploit shoreline and water-surface geometry for GPS-denied autonomous surface vessel localization. The first fra...

Derek Benham, Joshua G. Mangelson · 0 citations
Preprint Sep 2026

STARS-GS: Structure-Aware Regularized Gaussian Splatting for Large-Scale Aerial Surface Reconstruction

Large-scale 3D surface reconstruction from aerial imagery is fundamental to geospatial mapping and urban modeling. Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated considerable potential for this task. However, existing methods still face three major challenges in large and complex scenes: scene partit...

Bo-Cheng Li, Wen-Juan Zhang, Jie-Pan-Dong-Xu Han et al. · 0 citations
Preprint Aug 2026

Ground-to-Satellite Localization in Unconstrained Image Collections for 3D Scene Reconstruction

Ground image localization with respect to satellite imagery is a key enabler for metrically-accurate, geo-localized 3D scene reconstruction from unconstrained image collections. Existing cross-view localization methods have strict requirements such as panoramic imagery or known initial locations, limiting their applica...

A. Daruna, Ben Southall, Niluthpol Chowdhury Mithun et al. · 0 citations
Preprint Aug 2026

Geospatial-Prior Guidance for 3D Semantic Scene Completion

GeoScene is a geospatially guided framework that jointly uses satellite imagery and structured OpenStreetMap cues as soft priors for 3D semantic scene completion and consistently improves both geometric and semantic completion under the geospatial-prior-assisted setting.

Meng Wang, Shougao Zhang, Wenzhe He et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.