CORE introduces a decorrelated feature alignment module to directly align heterogeneous features into a unified representation space, which retains their semantic information, and formulates unified TAD as an in-context reconstruction problem, eliminating the need for labeled or synthesized anomalies.
Abstract
Tabular anomaly detection (TAD), which focuses on identifying abnormal samples that deviate from the majority in tabular data, has received growing attention. Recently, there has been an emerging trend towards unified TAD, which seeks to detect anomalies across different datasets using a single generalizable model. In unified TAD, aligning heterogeneous data remains challenging. While existing methods often rely on distance-based unified feature construction, they may obscure the semantics of the original features. Moreover, existing approaches typically formulate anomaly detection as a binary classification task, which may overlook diverse anomaly patterns from various datasets and be misled by unrepresentative synthetic anomalies. To address these challenges, we propose an in-COntext REconstruction approach for unified TAD (CORE for short). It introduces a decorrelated feature alignment module to directly align heterogeneous features into a unified representation space, which retains their semantic information. Meanwhile, CORE formulates unified TAD as an in-context reconstruction problem, eliminating the need for labeled or synthesized anomalies. Specifically, the in-context reconstruction module reconstructs each sample by leveraging contextual normal samples to capture dataset-specific distributions, such that reconstruction errors reflect its deviation from normality, facilitating unified TAD on arbitrary unseen datasets.
LLM-Detector is proposed, a framework that utilizes the in-context learning capacity of LLMs for structured, prompt-conditioned scoring synthesis, enabling LLMs to derive anomaly detection logic from structured normal-state knowledge.
Tu Nguyen, Dang Nguyen, Thuc Duy Le et al.· 0 citations
Zero-shot anomaly detection (ZSAD) aims to identify and localize anomalies in previously unseen target domains without accessing any target-domain training data, which is crucial under privacy, security, or proprietary constraints. However, existing ZSAD methods often struggle to generalize across domains, as they are tightly coupled to specific object categories or rely on fragmented designs that fail to capture both semantic consistency and structural abnormality. In this paper, we propose UAD, a unified framework that addresses ZSAD from a holistic perspective by jointly modeling semantic regularity and anomaly-aware representations. The key insight of UAD is that effective ZSAD requires aligning multi-level semantic understanding with fine-grained structural cues, rather than relying solely on object-centric semantics or local appearance statistics. To this end, UAD organizes image representations into coherent semantic contexts and identifies anomalies as deviations from both local structural patterns and high-level semantic consistency. Furthermore, we enhance cross-domain robustness by improving semantic supervision and diversity through prompt concatenation and intensity-guided anomaly synthesis, enabling UAD to better generalize to unseen anomaly types and domains. Extensive experiments on 17 real-world anomaly detection datasets show that UAD achieves superior zero-shot performance of detecting and segmenting anomalies in datasets of highly diverse class semantics from various defect inspection and medical imaging domains. Our works are available at https://github.com/hanli6688/UAD
Yuqing Zhao, Min Meng, Jigang Wu et al.· IEEE Transactions on Image P...· 0 citations
Visual industrial anomaly detection has evolved from one-class modeling to more challenging multi-class settings, where diverse categories and complex visual patterns must be jointly handled. Existing approaches often assume that anomalies lie far from normal samples in feature or spatial space. However, this assumption frequently fails due to two key issues: cross-class semantic confusion, where normal structures of one category are misclassified as anomalies in another, and pixel similarity failure, where anomalous regions visually blend into normal backgrounds. To address these challenges, we propose RPGAD (Region-Prompt Guided Anomaly Detection), an information-theoretic framework that models anomalies as semantic predictive instability, reflected in the joint responses of dual paths. RPGAD integrates two components: 1) DPENet (Dual-Path regional Energy evaluation Network), which compares region-level responses across normal-only and mixed paths through an entropy-guided energy formulation to generate robust region prompts; and 2) RDNet (Reverse Distillation Network), which selectively reconstructs prompted regions and employs a Prototype-Contrastive Optimal Transport (PCOT) loss to enhance inter-class separability and local feature aggregation. Experiments on five anomaly detection benchmarks - MVTecAD, VisA, BTAD, MPDD, and Real-IAD - demonstrate the effectiveness of RPGAD. At $256 \times 256$ resolution, RPGAD achieves strong overall performance, with mAD of 87.8%, 80.3%, 85.2%, 86.2%, and 77.7% on five benchmarks, and pixel-level AP and F1-max gains of up to 12.2 and 10.4 points over strong baselines. These results confirm that RPGAD provides accurate and robust multi-class anomaly detection and localization in complex visual scenarios.
Industrial anomaly detection aims to identify and localize defective regions without relying on exhaustive annotations of all possible defect types. Although recent unsupervised methods have achieved strong performance, most are primarily designed for single-class settings and often struggle in multi-class scenarios, where diverse normal patterns may lead to over-generalization and reduce the discriminative capability between normal and anomalous regions. In this paper, we propose SwinAD, a reconstruction-based framework for multi-class unsupervised anomaly detection that leverages a frozen pretrained Swin Transformer V2 encoder and a feature diversity-preserving reconstruction decoder. The hierarchical encoder provides semantically rich multi-scale features, while stage-wise bottleneck modules with dropout prevent trivial identity mapping and encourage robust reconstruction of normal patterns. To further improve localization, we introduce a feature diversity-preserving reconstruction framework that maintains complementary reconstruction hypotheses instead of relying on a single decoding branch. The discrepancies between encoder features and the two reconstructed features are then aggregated across multiple scales to produce the final anomaly map. Experiments conducted on three industrial anomaly detection benchmarks, including MVTec AD, VisA, and Real-IAD, demonstrate that SwinAD achieves competitive image-level performance and strong pixel-level localization accuracy, with particularly notable improvements in pixel-level AP and 1 on MVTec AD. These results indicate that combining hierarchical Swin features with diverse multi-scale reconstruction substantially improve pixel-level localization in multi-class unsupervised anomaly setting.
Huong Ninh, Chien Thai, M. Trang et al.· 0 citations
A novel framework, Generate and Filter graph learning for Graph Anomaly Detection (GFGAD), which generates a diverse set of synthetic anomalies with enriched feature and structural information to balance the data distribution and significantly outperforms state-of-the-art baselines.
Mengyu Li, Yonghao Liu, Ximing Li et al.· IEEE Transactions on Pattern...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.