Skip to content

A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees

Sep 2026 · 0 citations · 32 references
Computer Science

TL;DR

This work revisits model completeness: how well the concepts can reproduce the model's outputs, and shows that model incompleteness of the concepts can be bounded by the autoencoder's reconstruction error.

Abstract

Concept-based explanations describe neural network predictions through human-understandable properties of inputs called concepts. The field encompasses approaches that differ in how they define and represent concepts and connect them to model predictions. We introduce a theoretical framework that describes these approaches in a common mathematical language and supports a shared analysis of their properties. For concept discovery, which identifies concepts automatically within a latent space of a trained model, we employ a concept autoencoder view. An encoder extracts concept representations from the model's latent space, and a decoder uses them to reconstruct the original latent representation. The autoencoder's reconstruction error measures how accurately its decoder recovers the original latent representation. We revisit model completeness: how well the concepts can reproduce the model's outputs. We show that model incompleteness of the concepts can be bounded by the autoencoder's reconstruction error. The autoencoder view also provides a common way to define individual concept attributions, which measure each concept's contribution to a prediction. We establish when these attributions sum to the model's prediction, and bound the discrepancy otherwise, thus providing attribution completeness guarantees.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

A Unifying Perspective on Language Model Representations: From Filler-Role Structure to Mechanistic Interpretability

This work proposes using Tensor Product Representations (TPRs) as a unifying hypothesis, and shows that TPRs can unify several prior interpretability methods: additive analogies, linear probing, sparse autoencoders, and activation patching.

Enshang Zhang, R. Thomas McCoy · 0 citations
#machine learning Preprint Oct 2026

Revisiting Explainable AI through Model-Independent Concept Dictionaries

Modern applications of AI rely on increasingly complex models. Explainable AI (XAI) has emerged as a set of techniques aimed at improving model transparency. However, existing XAI methods typically assume input features to be inherently interpretable, or they rely on intermediate internal abstractions that are difficul...

T. Schnake, Doreen Schöppenthau, Alexander Meyer et al. · 0 citations
#machine learning Preprint Sep 2026

ProToMEx: Rapid, Interpretable Explanations via Structured Representations

It is demonstrated empirically that ProToMEx not only produces explanations of comparable fidelity to popular methods like SHAP and LIME but also drastically reduces the amortised computational cost of generating local explanations, making it highly suitable for real-time applications.

A. Georgara, Adarsh Valoor, Sarvapali D. Ramchurn · 0 citations
#artificial intelligence Preprint Oct 2026

DisParQ: Self-Supervised Part Concepts for Interpretable Vision Foundation Models

Concept-based vision models represent images through an intermediate layer of human-inspectable concepts, so what a model relies on can be traced to those concepts. However, those models are often limited to fixed categories or depend on language to define their concepts. We introduce DisParQ (Discrete Parts with Quant...

Adam Pardyl, Siddhartha Gairola, Sukrut Rao et al. · 0 citations
#artificial intelligence Preprint Oct 2026

When Are Concept Bottleneck Model Explanations Faithful and Compact?

Concept bottleneck models (CBMs) are neural classifiers that allow to explain their decisions via high-level concepts, potentially enabling understanding, steering and debugging. However, their explanations are often derived heuristically. Building on formal explainability, we argue they should also be faithful, i.e.,...

Stefano Teso, E. Marconato, Steve Azzolin et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Operational Abstractions of Neural Network Concepts via Topological Representations

Concept-based methods provide a semantic level for interpreting and manipulating learned representations, but existing editing approaches are typically specialized to particular interventions and do not provide a common and editable representation of concept organization. To achieve this, we introduce Topological Conce...

Mathieu Pont, Christoph Garth · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.