Jul 2026· 2026 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW)· pp. 278-289· 0 citations· 73 references
Computer Science
Abstract
Sparse Autoencoders (SAEs) offer a promising unsupervised interpretability approach for extracting human-interpretable concepts from large language models. Yet, their use in the security domain remains underexplored. Security-related classification tasks typically rely on smaller models than those commonly studied with SAEs. In this paper, we examine how SAEs can be used to interpret classification models fine-tuned for security tasks. We apply an interpretability framework that combines established techniques for foundation models to generate concept explanations, focusing on two widely studied problems in safety and security: hate speech and deepfake detection. We demonstrate its ability to produce meaningful concept explanations while identifying critical challenges for the effective deployment of SAEs in security contexts. Our findings suggest that while SAEs offer a promising unsupervised technique for generating concept explanations, addressing the identified challenges is necessary for their useful application in security interpretability.
XSec is introduced, a self-explainable deep architecture developed for security applications that produces deterministic explanations for a fixed trained model and input and substantially reduces explanation latency compared with approximation-based and perturbation-based post-hoc methods.
CON decomposition is introduced, which quantifies how much of a layer's variance each concept explains given all other concepts and the outcome, and how much none of them explains, yielding layer-comparable, calibrated scores that suppress false positives.
R. Rane, Marco Simnacher, Manuel Pfeuffer et al.· 0 citations
A systematic study of how pruning affects SAE behavior is presented and theoretically shows that, for a fixed SAE, its impact is governed by perturbation energy, a covariance-weighted norm.
Suchit Gupte, Xue-Ru Zhang, M. Khalili· 0 citations
Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from co...
Weihang Meng, Hongzhu Guo, Yi Jing et al.· 0 citations
Feature-robust Augmentation is introduced, which comprises diversified degradation-aware augmentation strategies, and a supervised contrastive learning pattern paired with a mean-teacher architecture that stabilizes features against augmentations through consistency constraints that wins the first place in ACM Multimed...
Zhu Xu, Jia-Qi Tang, Po-Kai Chen et al.· 0 citations
This work explores CXAI use cases in multi-label classification by training two DNNs, VGG16 and ResNet50, on the 20 most annotated labels in the MS-COCO dataset and demonstrates the potential of CXAI to enhance the understanding of model generalizability and to diagnose bias instigated by the dataset.
Haadia Amjad, Ronald Tetzlaff· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.