Bridging Optical Appearance and SAR Scattering With Vision-Language Prototypes for Zero-Shot Target Recognition
Abstract
Supervised synthetic aperture radar (SAR) automatic target recognition (ATR) methods rely on a closed-set assumption and struggle to recognize unseen target categories without labeled SAR samples. Zero-shot learning (ZSL) offers a promising solution by transferring knowledge from seen classes and external semantic priors. However, linguistic prototypes are weakly related to SAR scattering characteristics, and vision-language model (VLM) priors learned from optical image–text pairs cannot be directly transferred to SAR representations. To bridge optical appearance and SAR scattering, this letter proposes a prototype alignment framework (VLPA) via vision-language prototypes. The proposed method constructs vision-language prototypes by injecting text-guided optical residuals into textual semantic anchors, producing visually grounded category representations for both seen and unseen classes. A SAR-guided prototype alignment strategy is further designed to align SAR scattering features with the prototype space while preserving cross-modal relational consistency. Experiments show that VLPA consistently outperforms existing methods, achieving maximum Top-1 and harmonic mean gains of 1.91% and 1.53% on ATRNet-STAR, and 1.67% and 1.69% on FUSARShip with FGSC-23, respectively. These results demonstrate that vision-language prototypes provide an effective bridge for improving zero-shot SAR target understanding.