Jul 2026· The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences· 0 citations· 4 references
TL;DR
A confidence-weighted Sim (3) registration algorithm that utilizes a learned confidence mask to filter out unreliable points in dense street-level reconstructions and restores absolute scale and global coordinates without requiring high-grade GNSS/INS or expensive on-board LiDAR systems.
Abstract
Abstract. We present a framework that integrates ground-level imagery with Airborne Laser Scanning (ALS) point clouds. While the Visual Geometry Grounded Transformer (VGGT) enables dense geometry estimation from uncalibrated images, its application is limited by non-metric results and high GPU requirements. By leveraging publicly available, georeferenced ALS point clouds as an external metric constraint, our system restores absolute scale and global coordinates without requiring high-grade GNSS/INS or expensive on-board LiDAR systems. We introduce a confidence-weighted Sim (3) registration algorithm that utilizes a learned confidence mask to filter out unreliable points in dense street-level reconstructions. Experimental evaluations conducted on large-scale urban datasets demonstrate the average check point errors of 0.77 meters in Hong Kong dataset and 0.69 meters in Wuhan dataset, showing great potentials of feed-forward models in large-scale outdoor dense mapping.
Abstract. Semantic classification is a fundamental step in Mobile Laser Scanning (MLS) point clouds processing, and remains a non-trivial task. In this work, we propose a classification framework based on a 3D Sparse Convolutional Neural Network (SparseCNN) for efficient processing of large-scale MLS data. A coarse-to-fine two-stage pipeline is introduced, where an essential model performs a classification for the entire scene, followed by a refinement stage for detailed ground-surface classes. To enhance robustness under diverse acquisition conditions, both point-wise and scene-wise data augmentation strategies are employed during the training, including rotation, jittering, density perturbation, noise injection, and patch swapping. To account for environmental and sensor variations, wavelength-specific models are trained for both urban and highway scenes. Experimental results on urban and highway datasets demonstrate strong performance, achieving over 90% accuracy for major classes, while ablation studies show that radiometric features are critical for distinguishing material dependent classes, such as traffic signs, and that the proposed augmentation strategies improve performance for challenging object categories, such as pedestrian, which is dynamic and structurally ambiguous.
Nan-Feng Li, H. Teufelsbauer, F. Pöppl et al.· The International Archives o...· 0 citations
Point cloud fusion is crucial in geospatial analysis, combining data from multiple sources (e.g, LiDAR and photogrammetry) to provide a more complete and accurate environmental representation. However, integrating airborne hybrid sensors or cross-source point clouds remains challenging due to variations in geometric accuracy, data precision, gaps, and sensor attributes. Despite recent advancements, these challenges remain and are among the most demanding aspects in geospatial data processing for remote sensing applications. We propose a new point cloud fusion algorithm that leverages local plane constraints to achieve advanced semantic consistency. The proposed method dynamically fits local planes to the target point clouds, enabling robust alignment of source points to these planes. Evaluation on two real-world datasets demonstrates significant gains in accuracy and preservation of geometric details. Our algorithm also improves the accuracy of downstream tasks such as semantic segmentation. In our experiment, the overall accuracy for the Dudelange dataset increases from 48.5% to 80.1%, and that for the Dublin dataset increases from 72.9% to 88.0%. While challenges persist with sparse and noisy datasets, experimental results highlight the effectiveness of the proposed method, offering valuable insights for maximizing the potential of cross-source point cloud data.
Shahoriar Parvaz, Félicia N. Teferle, Abdul Nurunnabi et al.· Remote Sensing· 1 citation
Existing point-based generative methods for outdoor scenes primarily focus on LiDAR-conditioned completion. During training, noisy point clouds are constructed by perturbing complete ground-truth scenes, whereas during inference, they are initialized by adding noise to duplicated partial scans. This train-inference mismatch inherits the sparsity and visibility bias of partial scans, leading to sparse distant regions and incomplete geometry in occluded areas. Moreover, the reliance on partial scans restricts generation when LiDAR observations are unavailable or replaced by layout cues. We present FPSGen, a flexible framework that constructs point sources independently of partial scans. FPSGen first predicts a bird's-eye-view (BEV) prior with density, height, and mask channels from the active cues. The density map is then sampled to form a BEV-supported point source, enabling both unconditional and conditioned initialization. A teacher-student approximate optimal transport scheme then uses teacher-predicted endpoints to learn a velocity field that induces straighter transport paths. By integrating BEV point source construction with path-straightening transport, FPSGen provides a unified framework for unconditional and flexible cue-conditioned scene generation. Extensive experiments show that FPSGen achieves state-of-the-art JSD and voxel IoU performance on SemanticKITTI completion while maintaining strong performance with a single point transport step. On KITTI-360 unconditional generation, it also achieves the best Coverage (COV) among the compared methods.
Abstract. Satellite imagery offers a distinct advantage in Earth observation by providing expansive coverage and enabling the monitoring of inaccessible regions without physical on-site intervention, serving as a significantly more cost-effective and scalable alternative to traditional aerial or ground-based surveys. The task of 3D reconstruction from multi-view satellite images has therefore been a pivotal point of research at the intersection of photogrammetry and remote sensing. Recently, novel-view synthesis techniques such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have accelerated the accuracy and speed of topographic modeling. Among these, Earth Observation Gaussian Splatting (EOGS) has emerged as a state-of-the-art approach by adapting 3DGS to handle the unique geometric and radiometric characteristics of satellite data, including Rational Polynomial Coefficients (RPCs) and varying solar conditions. However, the standard EOGS pipeline relies on stochastic initialization, where Gaussians are distributed uniformly within a volumetric bounding box, leading to high computational overhead and dependency on aggressive pruning that can inadvertently remove critical geometric features, particularly in areas with complex urban structures. To address these limitations, we propose Bundle-Adjusted Initialization for Earth Observation Gaussian Splatting, which leverages sparse point clouds from bundle adjustment as geometric priors for Gaussian initialization. Combined with an adaptive densification strategy, our method achieves faster convergence and improved DSM accuracy on the DFC2019 dataset compared to the EOGS baseline.
Jiyong Kim, Shuang Song, Rongjun Qin· The International Archives o...· 0 citations
Abstract. Recent handheld scanners increasingly integrate geometric (LiDAR-based) and visual (image-based) SLAM (Simultaneous Localization And Mapping), promising low-cost and flexible solutions for surveying tasks. This paper evaluates the accuracy of three such systems: the XGRIDS Lixel K1, the SHARE S20, and a Pix4D solution pairing an iPhone Pro with an Emlid Reach RX GNSS (Global Navigation Satellite System) antenna. We conducted experiments in two distinct environments: Scene 1, with continuous, high-quality RTK (Real-Time Kinematic) coverage, and Scene 2, which included an indoor trajectory resulting in a temporary loss of the RTK fix. Accuracy was validated against independent GNSS check points. In Scene 1, the Pix4D solution delivered survey-grade results, achieving a RMSE (Root Mean Square Error) below 3 cm in the X, Y , and Z directions. The XGRIDS and SHARE scanners yielded larger maximum errors, around 10 to 15 cm. In Scene 2, accuracy degraded; the Pix4D solution’s maximum error increased to approximately 12 cm , while the Share S20’s maximum error exceeded 25 cm. We conclude that while the fusion of visual and geometric SLAM is powerful, a stable RTK fix remains critical for achieving consistent surveygrade accuracy with current low-cost handheld scanners.
Christoph Strecha, Ryan Hughes, Davide A. Cucci et al.· The International Archives o...· 0 citations
Abstract. The emergence of LiDAR-equipped mobile devices has enabled low-cost 3D point cloud acquisition. However, accurate 3D model reconstruction using such devices remains confined to individual, locally captured scans. This limitation makes wide-area coverage infeasible without additional correction. Existing methods for wide-area mapping require specialized equipment or high-accuracy base maps, which poses significant barriers in terms of cost and data infrastructure availability. This study proposes a low-cost pipeline for constructing wide-area 3D maps using only LiDAR-equipped mobile devices and open data. The proposed approach automatically registers multiple locally captured point clouds and assigns absolute coordinates to each by referencing open geospatial datasets. A key feature of the method is the decomposition of 3D spatial information into horizontal and vertical components. These components are processed independently to reduce errors and improve computational efficiency. Validation experiments in Tondabayashi City, Osaka, Japan confirmed a horizontal RMSE of 1.75 m or less and a vertical RMSE of 0.10 m or less. Both values meet the accuracy requirements for 1:2,500-scale disaster prevention base maps. The vertical accuracy in particular exceeds the stricter standard required for flood hazard mapping applications. These results demonstrate that combining mobile devices with open data enables non-experts to construct high-accuracy 3D maps suitable for disaster risk management applications.
Ryosei Ueda, Daisuke Yoshida· The International Archives o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.