This work proposes ElainaCLIP, a prompt learning framework based on CLIP to incorporate vision-guided information into text prompt representations and introduce ElainaLoss to guide prompt learning through low-level semantic constraints, thereby enhancing the modeling of unstructured anomaly semantics.
Zero-shot industrial anomaly detection (ZIAD) aims to develop a unified model capable of directly identifying unseen anomaly categories in images without requiring reference samples. Recently, large-scale Vision-Language Models (VLMs) such as CLIP have shown great potential for solving this task. However, existing meth...
Tiyu Fang, Lin Zhang, Ran Song et al.· IEEE Transactions on Automat...· 0 citations
A unified zero-shot multimodal anomaly detection framework GRASP is proposed, with substantial improvements of 3.0 points in I-AUROC and 2.7 points in AUPRO compared to the SOTA method.
Zhong-Bin Sun, Yu-Ze Cui, Yong-Yi Zhou· Proceedings of the Thirty-Fi...· 0 citations
Zero-shot anomaly detection (ZSAD) aims to identify anomalies in target datasets without accessing their samples. Although CLIP and other large-scale vision language models show strong generalization, their potential for multi-level feature extraction in ZSAD remains underexplored. To address this, we propose a Multi...
Jianfeng Qiu, Junfa Li, Juan Xie et al.· Complex & Intelligent Sy...· 0 citations
FreqPrompt-AD achieves the strongest average performance among the compared zero-shot VLM-based detectors in the authors' local full-scale evaluation, improving average image-level and pixel-level AUROC over DLVP-CLIP by 2.15 percentage points.
Siqi Qiao, Chang Liu· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.