Skip to content

Towards Explaining Classification Models in Security with Sparse Autoencoders

Jul 2026 · 2026 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW) · pp. 278-289 · 0 citations · 73 references
Computer Science

Abstract

Sparse Autoencoders (SAEs) offer a promising unsupervised interpretability approach for extracting human-interpretable concepts from large language models. Yet, their use in the security domain remains underexplored. Security-related classification tasks typically rely on smaller models than those commonly studied with SAEs. In this paper, we examine how SAEs can be used to interpret classification models fine-tuned for security tasks. We apply an interpretability framework that combines established techniques for foundation models to generate concept explanations, focusing on two widely studied problems in safety and security: hate speech and deepfake detection. We demonstrate its ability to produce meaningful concept explanations while identifying critical challenges for the effective deployment of SAEs in security contexts. Our findings suggest that while SAEs offer a promising unsupervised technique for generating concept explanations, addressing the identified challenges is necessary for their useful application in security interpretability.

View source

Similar papers

Preprint Aug 2026

A Self-Explainable Deep Architecture for Security Applications

XSec is introduced, a self-explainable deep architecture developed for security applications that produces deterministic explanations for a fixed trained model and input and substantially reduces explanation latency compared with approximation-based and perturbation-based post-hoc methods.

Ananth Shreekumar, Jyun-Jhu Syu, Muslum Ozgur Ozmen et al. · 0 citations
#artificial intelligence Preprint Aug 2026

ICON Decomposition: Auditing Deep Neural Networks with Multivariate Variance-based Concept-level Explanations

CON decomposition is introduced, which quantifies how much of a layer's variance each concept explains given all other concepts and the outcome, and how much none of them explains, yielding layer-comparable, calibrated scores that suppress false positives.

R. Rane, Marco Simnacher, Manuel Pfeuffer et al. · 0 citations
Preprint Aug 2026

SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from co...

Weihang Meng, Hongzhu Guo, Yi Jing et al. · 0 citations
Preprint Aug 2026

Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization

Feature-robust Augmentation is introduced, which comprises diversified degradation-aware augmentation strategies, and a supervised contrastive learning pattern paired with a mean-teacher architecture that stabilizes features against augmentations through consistency constraints that wins the first place in ACM Multimed...

Zhu Xu, Jia-Qi Tang, Po-Kai Chen et al. · 0 citations
Preprint Aug 2026

Identifying Confusion Trends in Concept-based XAI for Multi-Label Classification

This work explores CXAI use cases in multi-label classification by training two DNNs, VGG16 and ResNet50, on the 20 most annotated labels in the MS-COCO dataset and demonstrates the potential of CXAI to enhance the understanding of model generalizability and to diagnose bias instigated by the dataset.

Haadia Amjad, Ronald Tetzlaff · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.