Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for scene understanding under diverse conditions. However, differences in optics, resolution, and mounting often limit practical systems to global or image-center alignment. After tokenization, parallax, platform motion, and lens distortion can shift corresponding patch centers across modalities, weakening the spatial correspondence assumed by dense contrastive learning and cross-modal fusion. We propose GAAT (Geometry-Aware Alignment Transformer), an alignment-first pretrained model that estimates local correspondence reliability before cross-modal interaction. GAAT introduces syncPATC, which learns patch-center consistency under synchronized view transformations without correspondence annotations. It emits geometric priors, including token and query confidence, query centers, and sub-token offsets, that identify reliable local anchors across residual misalignment. Guided by these priors, MG-Sparse-MMA performs query-mediated sparse fusion over top-K_s reliable regions, replacing dense all-patch interaction with geometry-calibrated local updates. RA-QCGCL aligns pretraining supervision with this sparse query bottleneck through reliable patch-to-patch, patch-to-query, and query-to-query contrastive branches. We introduce UAVMeta and StateBench, which provide four acquisition-state scores derived from platform telemetry and image statistics: camera reliability, observation scale, viewpoint stability, and flight maneuver complexity. Extensive experiments across six downstream tasks demonstrate consistently superior transfer performance, establishing GAAT as a state-of-the-art multimodal foundation model for UAV perception. StateBench further enables a systematic diagnosis of real-world acquisition conditions.
Jing-Pu Yang, Deming Tang, Yi-Lin Sun et al.· 1 citation
Underwater acoustic target recognition (UATR) is challenging due to the complex, multi-scale physical characteristics of marine targets and the strict computational limits of edge platforms like unmanned surface vehicles. To navigate the severe interference of underwater environments, existing methods increasingly rely on heavyweight architectures to achieve high recognition accuracy. However, the massive computational overhead of these models is fundamentally at odds with the restricted power and processing capabilities of practical deployment platforms. To resolve this conflict between performance and deployability, we propose LHK-Net, a lightweight Heterogeneous Kernel Network. By integrating a Heterogeneous Kernel Pyramid with Residual Depthwise Separable Convolutions, LHK-Net dynamically captures multi-scale acoustic features, from macroscopic steady-state harmonics to localized transient impulses, while compressing the model size to merely 0.82 M parameters. Additionally, a dual-domain Time–Frequency Attention module and an Adaptive SK-Fusion mechanism are incorporated for robust noise suppression. Experiments on the DeepShip dataset demonstrate that LHK-Net achieves state-of-the-art accuracy, outperforming heavyweight models at real-time speeds. Extensive visual analyses further validate that the network possesses strong physical interpretability, effectively aligning its internal feature representations with the intrinsic acoustic properties of the targets.
Yilling Sun, Menghao Fan, Haonan Wei et al.· Journal of Marine Science an...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.