Abstract. Airborne laser scanning (ALS) point clouds are widely used for large-scale 3D scene understanding, but acquiring dense ALS data remains costly and sparse observations often exhibit incomplete vertical structures and uneven sampling. Existing diffusion-based point cloud upsampling methods have shown promise, yet conditioning only on sparse point coordinates often leads to noisy surfaces, boundary artifacts, and limited structural recovery in under-observed regions. In this work, we propose the Appearance-aware Scaling Diffusion Model (ASDM), a conditional diffusion framework for scene-level ALS point cloud upsampling that incorporates multi-view projected depth-image cues derived directly from the input point cloud. Specifically, sparse ALS scenes are rendered from multiple virtual viewpoints to generate projected depth images, which provide complementary structural information for guiding the denoising process. These projected-image features are fused with sparse point features to improve geometric fidelity and scene-level consistency. For training and evaluation, we construct realistic sparse–dense scene pairs from aerial LiDAR data derived from the YUTO Semantic dataset using a region-disjoint split over 13 survey areas. Experiments under the ×4 upsampling setting show that ASDM outperforms recent diffusion-based baselines, achieving the best overall performance in Chamfer Distance (0.5643), JSD-3D (0.6688), F1-score (75.67), and voxelized IoU at 1m and 2m resolutions. These results demonstrate that projected-image conditioning is an effective strategy for robust airborne LiDAR scene densification.
Sunghwan Yoo, Gunho Sohn· The International Archives o...· 0 citations
Abstract. Surface-laid unexploded ordnance (UXO) and landmines constitute a critical humanitarian crisis. While unmanned aerial vehicles (UAVs) provide a scalable remote sensing solution, detecting modern, non-metallic explosive devices in cluttered environments remains a profound Camouflaged Object Detection (COD) challenge. Traditional optical sensors frequently suffer from foreground-background confusion when a target’s texture mimics its surroundings. To overcome these physical bottlenecks, we introduce XPol- Net, a novel multimodal architecture synergizing the semantic reasoning of Vision Transformers with the deterministic physics of polarimetric imaging. Built on a hierarchical PVTv2 backbone, XPol-Net utilizes a progressive Dual Cross-Attention Strategy for effective modality fusion. In early stages, Channel Cross-Attention (CCA) filters material-specific Degree of Linear Polarization (DoLP) cues to suppress background clutter. In deeper stages, Spatial Cross-Attention (SCA) dynamically aligns high-level RGB semantics with strict structural boundaries. To enhance robustness and prevent modality collapse, we deploy a multi-task auxiliary learning framework that reconstructs the continuous Angle of Linear Polarization (AoLP) map. On the PCOD benchmark, XPol-Net achieves state-of-the-art results in global structural alignment (Eϕ of 0.980 and 0.984 at 352 × 352 and 704 × 704, respectively). While minor trade-offs are observed in localized metrics such as Sα or Fβ, XPol-Net remains highly competitive, consistently delivering superior results in Eϕ and MAE. By prioritizing structural recall over localized strictness, XPol-Net ensures the complete discovery of concealed targets, establishing a reliable, physics-aware foundation for humanitarian demining operations.
Youssef Korny, Sunghwan Yoo, Gunho Sohn· The International Archives o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.