HD-CLIP: Hierarchical Dynamic Prompting and Decoupled Learning for Zero-Shot Anomaly Detection
The potential of vision-language models (VLMs) such as CLIP for zero-shot anomaly detection (ZSAD) is constrained by an inherent semantic-localization dichotomy. While CLIP’s global features excel at image-level classification, they lack the spatial sensitivity required for pixel-level segmentation. Existing approaches attempt to alleviate this issue through prompt optimization, which introduces a trade-off between global semantic discrimination and local localization accuracy, thereby limiting cross-domain generalization capability. To resolve this, we introduce HD-CLIP, a framework that decouples these competing objectives. HD-CLIP employs dedicated pathways for classification and localization, guided by a hierarchical and dynamic prompting mechanism that provides multilevel, content-adaptive cues. A novel localization-distilled supervision (LDS) loss then unifies these pathways by creating a probabilistic bridge between spatial evidence and semantic judgment. Extensive experiments on 14 challenging datasets achieve state-of-the-art performance, confirming that HD-CLIP effectively bridges the semantic-localization divide and advances the potential of VLMs for robust ZSAD applications.