Skip to content

HD-CLIP: Hierarchical Dynamic Prompting and Decoupled Learning for Zero-Shot Anomaly Detection

2026 · IEEE Transactions on Instrumentation and Measurement · Vol 75, pp. 5014411-5014411 · 0 citations · 44 references

Abstract

The potential of vision-language models (VLMs) such as CLIP for zero-shot anomaly detection (ZSAD) is constrained by an inherent semantic-localization dichotomy. While CLIP’s global features excel at image-level classification, they lack the spatial sensitivity required for pixel-level segmentation. Existing approaches attempt to alleviate this issue through prompt optimization, which introduces a trade-off between global semantic discrimination and local localization accuracy, thereby limiting cross-domain generalization capability. To resolve this, we introduce HD-CLIP, a framework that decouples these competing objectives. HD-CLIP employs dedicated pathways for classification and localization, guided by a hierarchical and dynamic prompting mechanism that provides multilevel, content-adaptive cues. A novel localization-distilled supervision (LDS) loss then unifies these pathways by creating a probabilistic bridge between spatial evidence and semantic judgment. Extensive experiments on 14 challenging datasets achieve state-of-the-art performance, confirming that HD-CLIP effectively bridges the semantic-localization divide and advances the potential of VLMs for robust ZSAD applications.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.