Aug 2026· Journal of King Saud University: Computer and Information Sciences· Vol 38· 0 citations· 53 references
TL;DR
FreqPrompt-AD achieves the strongest average performance among the compared zero-shot VLM-based detectors in the authors' local full-scale evaluation, improving average image-level and pixel-level AUROC over DLVP-CLIP by 2.15 percentage points.
Abstract
Zero-shot industrial anomaly detection aims to identify and segment defects in unseen categories without target-specific model fitting. CLIP-style vision-language models (VLMs) provide useful semantic transferability, but their global image-text alignment often emphasizes object-level semantics and weakens the response to local high-frequency defects. To address this granularity mismatch, we propose FreqPrompt-AD, a training-free framework that constructs local-frequency semantic evidence with a frozen VLM. Frequency-aware prompt interaction (FPI) converts high-frequency residuals into semantic-gated spatial guidance for patch-text anomaly evidence; local semantic calibration (LSC) removes object-level semantic dominance and builds image-adaptive normality prototypes; and text-anchored decision (TAD) integrates semantic and prototype-deviation evidence into image-level anomaly scores and dense anomaly maps. We further evaluate FreqPrompt-AD-FT, an adapter-enhanced variant in which only a lightweight adapter is optimized while the VLM backbone and text encoder remain frozen. Full-scale experiments on MVTec AD, VisA, BTAD, and MPDD show that FreqPrompt-AD achieves the strongest average performance among the compared zero-shot VLM-based detectors in our local full-scale evaluation, improving average image-level and pixel-level AUROC over DLVP-CLIP by 2.05 and 2.15 percentage points, respectively. Additional frequency-prior comparisons, robustness evaluations, feature interaction analysis, and qualitative results validate the effectiveness and practical applicability of the proposed local-frequency semantic evidence formulation.
Zero-shot anomaly detection (ZSAD) aims to identify anomalies in unseen domains, a setting that is particularly critical for industrial and medical applications where domain shifts are prevalent. However, most CLIP-based ZSAD methods anchor semantics solely on the text modality, making performance highly sensitive to p...
CLIP-based anomaly detectors have markedly advanced training-free and zero-shot industrial anomaly detection and localization, yet their predictions remain dominated by patch-wise vision–language similarity or anomaly-aware feature scoring. This formulation is intrinsically limited for logical anomalies, in which every...
NOVA is proposed, a training-free ZS-VAD framework that strengthens the normal side at both linguistic and visual levels, and introduces Normality-Aware Prompt Construction (NA), which excludes anomaly-adjacent verbs and biases normal descriptions toward static, low-motion scenes.
Wei-Chih Yin, Yu-ching Kao, Cheng-Kuan Lin et al.· 0 citations
Vision Foundation Models (VFMs) provide transferable patch representations for few-shot industrial anomaly detection, but their attention computation is typically inherited from pretraining objectives centered on semantic aggregation. This creates a potential mismatch: token relations that support semantic recognition...
Xiao-Yu Yang, Qi-Xing Wu, Hui Zhao et al.· 0 citations
Few-shot anomaly detection (FSAD) has recently benefited from vision-language models such as CLIP, which enable anomaly de?tection by aligning visual features with text descriptions of normal and abnormal states. However, existing methods typically rely on static text prompts that are applied uniformly across the entir...
Wen-Yang Liu, Tianyi Liu, Dongshuo Zhang et al.· 0 citations
This work proposes ElainaCLIP, a prompt learning framework based on CLIP to incorporate vision-guided information into text prompt representations and introduce ElainaLoss to guide prompt learning through low-level semantic constraints, thereby enhancing the modeling of unstructured anomaly semantics.
Chun-Lei Wu, Yong-Hao Wang, Bai Qiao et al.· Multimedia Systems· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.