Vision-guided multi-image prompt learning for zero and few-shot pipeline anomaly detection
This work proposes ElainaCLIP, a prompt learning framework based on CLIP to incorporate vision-guided information into text prompt representations and introduce ElainaLoss to guide prompt learning through low-level semantic constraints, thereby enhancing the modeling of unstructured anomaly semantics.