This work proposes a complementary strategy that leverages three-dimensional human models reconstructed via photogrammetric techniques that offers a practical and effective enhancement for UAV-based human detection pipelines where timely, reliable identification is critical.
Abstract
Recently, deep learning has enabled unmanned aerial vehicles (UAVs) to detect human bodies in aerial imagery, which is of particular importance in post-disaster situations such as floods and storms. Yet progress in this domain remains constrained by a familiar obstacle: the shortage of annotated training data. Neural networks, while powerful, are highly sensitive to data volume and diversity. Existing augmentation strategies help reduce this gap but typically introduce only incremental novelty, especially with respect to viewpoint variation, thereby limiting dataset richness. In this work, we propose a complementary strategy that leverages three-dimensional human models reconstructed via photogrammetric techniques. By situating these models within a controlled rendering environment, we generate synthetic imagery across a broad range of elevations and camera angles—perspectives that are rarely captured in conventional UAV datasets. These additions are designed to increase both the variability and the resilience of the training corpus. To evaluate the contribution of this approach, a custom CNN deep convolutional neural classifier was trained and benchmarked on a UAV human vs. non-human patch dataset of 4000 baseline images (128 × 128 px; 2800 train, 600 validation, 600 test), expanded with 3000 photogrammetry-derived synthetic patches (balanced by class) to 7000 total images for the 3DG setting. The primary metric was classification accuracy on the held-out test set, consistent with patch-level evaluation practice; detection-style metrics such as AP/IoU were not applicable to this binary classification protocol. Averaged over five independent training runs, the proposed augmentation improved classification accuracy by 3.02 percentage points over the baseline (88.06 ± 0.97% → 91.08 ± 1.03%), with consistent gains in precision, recall, and F1-score. When combined with standard augmentations (rotation, translation, scaling, flipping), accuracy reached 95.21 ± 0.61%, a gain of 7.15 percentage points over the baseline. These results suggest that photogrammetry-based augmentation offers a practical and effective enhancement for UAV-based human detection pipelines where timely, reliable identification is critical.
A hybrid framework that decouples detection from damage assessment is proposed, combining the precision of CV models with the reasoning power of LVLMs, and the best combination under this framework accurately counts intact, partially damaged and completely destroyed buildings.
H. Ung, Guillaume Habault, Roberto Legaspi et al.· 0 citations
Traditional feature-based image stitching methods depend heavily on the quality of feature matching, which leads to suboptimal performance when applied to drone remote sensing images with significant differences in viewpoint and depth of field. Concurrently, supervised learning paradigms have proven infeasible due to t...
Wenpeng Zhang, Xiangyue Zhang, Huaguang Shi et al.· IEEE Geoscience and Remote S...· 0 citations
Flood disasters consistently cause massive damage every year, making rapid mapping of affected areas crucial for coordinating emergency aid. The use of unmanned aerial vehicles (UAVs) offers a practical solution to obtain high-resolution aerial imagery, but manually identifying flood areas from hundreds of images remai...
Fariida Aini, Muhammad Akrom, Gustina Alfa· JOURNAL OF APPLIED INFORMATI...· 0 citations
Vision Transformers (ViTs) perform well on clean aerial imagery but degrade sharply when deployed in post-earthquake UAV operations, where motion blur, dust haze, illumination variation, and sensor noise combine to produce what we term
post-earthquake visual drift
, a structured distributional shift that can render...
Jagan Murugesan, P. Vijitha, L. J. Vinita et al.· Frontiers in Artificial Inte...· 0 citations
Landmines continue to pose a severe threat to civilian populations in post-conflict regions, where conventional detection methods are often slow, risky, and ineffective against modern non-metallic mines. This paper presents an autonomous UAV-based landmine detection system that relies solely on thermal imaging and mach...
P. Santhoshini, K. Helenprabha, S. T. Jaibalaji et al.· International Conference on...· 0 citations
Ground-to-Air (G2A) drone detection in medium- and long-wave infrared (MWIR/LWIR) imagery is challenging due to reduced texture information, sensor noise, weak thermal contrast, and the scarcity of annotated data. This work investigates a synthetic-first training strategy that combines synthetic scene generation with f...
Tanel Liiv, Sander Soodla, Nzamba Bignoumba et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.