Skip to content

Vision-guided multi-image prompt learning for zero and few-shot pipeline anomaly detection

Sep 2026 · Multimedia Systems · Vol 32 · 0 citations · 58 references

TL;DR

This work proposes ElainaCLIP, a prompt learning framework based on CLIP to incorporate vision-guided information into text prompt representations and introduce ElainaLoss to guide prompt learning through low-level semantic constraints, thereby enhancing the modeling of unstructured anomaly semantics.

View source

Similar papers

2026

Cross-Modal Guidance Learning for Zero-Shot Industrial Anomaly Detection

Zero-shot industrial anomaly detection (ZIAD) aims to develop a unified model capable of directly identifying unseen anomaly categories in images without requiring reference samples. Recently, large-scale Vision-Language Models (VLMs) such as CLIP have shown great potential for solving this task. However, existing meth...

Tiyu Fang, Lin Zhang, Ran Song et al. · 0 citations
Open access Aug 2026

Multi-level visual-language models feature learning for generalizable anomaly detection

Zero-shot anomaly detection (ZSAD) aims to identify anomalies in target datasets without accessing their samples. Although CLIP and other large-scale vision language models show strong generalization, their potential for multi-level feature extraction in ZSAD remains underexplored. To address this, we propose a Multi...

Jianfeng Qiu, Junfa Li, Juan Xie et al. · 0 citations
Open access Aug 2026

FreqPrompt-AD: Frequency-guided local semantic prompting for zero-shot industrial anomaly detection and segmentation

FreqPrompt-AD achieves the strongest average performance among the compared zero-shot VLM-based detectors in the authors' local full-scale evaluation, improving average image-level and pixel-level AUROC over DLVP-CLIP by 2.15 percentage points.

Siqi Qiao, Chang Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.