Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

GeoThreat: Transferable Targeted Adversarial Attacks on Large Vision-Language Models for Remote Sensing Image Interpretation

Adversarial attacks against large vision-language models (LVLMs) serve as an effective means of assessing their robustness in cross-modal semantic understanding. Existing studies mainly focus on corrupting visual inputs to induce predefined erroneous responses in general vision-language tasks, whereas corresponding investigations in remote sensing fields remain largely underexplored. Compared with natural image understanding, remote sensing image interpretation requires joint reasoning over local discriminative cues and global scene context. This poses additional challenges to achieving transferable semantic manipulation toward specified responses under black-box settings. To tackle these challenges, we propose GeoThreat, a transferable targeted adversarial attack method against LVLMs for remote sensing image interpretation. Specifically, GeoThreat modulates adversarial representations in accordance with the target content at both conceptual and perceptual levels. The class tokens from surrogate image encoders are employed as conceptual representations, while perceptual representations are distilled from patch tokens of the adversarial example through collaborative importance estimation. Beyond merely rolling out attention scores across layers, we incorporate adversarial-target similarity gradients to more faithfully characterize the relevance of local visual cues to the intended semantic manipulation. The perceptual representations are then dynamically aligned with target patch tokens in a cross-attentive manner, facilitating the adaptation of local cues toward designated semantic details. Finally, adversarial perturbations are iteratively updated via ensemble-based joint optimization of conceptual calibration and perceptual adaptation. Extensive experiments across diverse LVLMs demonstrate the superiority of GeoThreat in both transferability and controllability.

Yimin Fu, Yuefeng Bai, Baicheng Pan et al. · 0 citations
2026

Reliable Tiny-Airborne-Object Detection via Prior-Guided Motion-Structure Verification

Reliable vision-based sensing of tiny-airborne objects is important for airspace surveillance and aerial monitoring. In practical scenarios, airborne objects are typically captured at long distances, occupying only a few pixels and exhibiting low signal-to-clutter ratios (SCRs) against complex backgrounds. These factors frequently cause missed detections and false alarms, thereby degrading reliable target detection and overall sensing reliability. Existing approaches seek to address these challenges by exploiting temporal cues across frames to stabilize weak object responses. However, they rarely impose explicit reliability-oriented verification on such cues, allowing clutter-induced spurious temporal responses to persist and undermine reliable detection. To address this issue, we propose motion-structure verification (MoSVer) for reliable learning, which introduces explicit joint verification of motion-derived temporal cues and structural cues to yield verified evidence for reliability-oriented supervision. Specifically, the motion-derived cue extraction (MDCE) module generates a motion support map to capture target-relevant motion evidence against the background. Meanwhile, the structure-constrained cue extraction (SCCE) module extracts a structural support map from regions exhibiting cross-frame structural consistency. Finally, the verification-guided matching (VGM) module integrates the two support maps to derive verified evidence, which is used as a reliability-aware matching prior to favor assignments supported by both cues. With RT-DETR as the base detector, MoSVer improves mAP ${}_{50:95}$ by 0.042 and 0.011 on airborne object tracking (AOT) and UAVSwarm, respectively, while reducing FP/frame@R = 0.50 by 17.8 % and 16.0 %. It also lowers MR@FP = 1 by 1.9 and 0.5 percentage points and decreases expected calibration error (ECE) by 0.025 and 0.009, providing quantitative evidence of improved detection reliability in clutter-dominated, low-SCR scenes.

Hengyang Zhao, Zhun-ga Liu, Changyuan Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.