Skip to content

Decomposing decisions: causal concept explanations for deep learning models

Jul 2026 · Journal of experimental and theoretical artificial intelligence (Print) · Vol 38, pp. 447 - 465 · 0 citations · 40 references
Computer Science

TL;DR

The Causal Concept Decomposer (CCD) is introduced, a three-stage framework for concept-driven causal explanation of object detectors that produces explanations that are both visually coherent and quantitatively more faithful than existing approaches.

Abstract

ABSTRACT The deployment of deep learning models in high-stakes applications such as autonomous driving is critically important, yet their black-box nature remains a fundamental barrier to trust and accountability. Existing explainability methods typically produce ambiguous, pixel-based heatmaps that capture correlation rather than establishing a causal link between high-level, human-interpretable concepts and model outputs. This paper introduces the Causal Concept Decomposer (CCD), a three-stage framework for concept-driven causal explanation of object detectors. CCD first employs semantic segmentation to isolate the target object, then applies Non-negative Matrix Factorization to discover constituent semantic parts, and finally uses Sobol sensitivity analysis to quantify the causal influence of each part on the detector’s decision. Evaluated on the MS COCO dataset, CCD produces explanations that are both visually coherent and quantitatively more faithful than existing approaches, achieving a Deletion AUC of 0.11 and an Insertion AUC of 0.91. By moving beyond correlational attribution towards principled causal analysis, this work represents an important step towards more trustworthy and reliable AI systems.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

ICON Decomposition: Auditing Deep Neural Networks with Multivariate Variance-based Concept-level Explanations

CON decomposition is introduced, which quantifies how much of a layer's variance each concept explains given all other concepts and the outcome, and how much none of them explains, yielding layer-comparable, calibrated scores that suppress false positives.

R. Rane, Marco Simnacher, Manuel Pfeuffer et al. · 0 citations
Preprint Aug 2026

Towards A Unified Information Bottleneck Framework for Time Series Explanations

Explaining deep learning models operating on time series data is crucial in various applications that require transparent and interpretable insights into model behavior. {Existing explanation methods generally fall into two categories: attribution-based explanations, which identify the temporal regions most responsible...

Xu Zheng, Zichuan Liu, Zhuomin Chen et al. · 0 citations
Preprint Aug 2026

Spatial Attention Noise Masking for Causally Sufficient Interpretability

We present a novel causal approach to interpretability for computer vision models that dynamically masks the input image prior to classification. The interpretability of deep learning predictions is critical in high-stakes fields such as medical imaging, security, and autonomous driving. Most interpretability methods a...

Benjamin Formby, Kuang-Ching Wang, D. H. Smith · 0 citations
Review Open access Jul 2026

CausalShift: A Modular, Plugin-Based Framework for Dataset Shift Handling in Machine Learning

CausalShift is proposed, a modular, plugin-based framework for end-to-end dataset shift handling that reduces the in-distribution to out-of-distribution accuracy gap, while remaining competitive on real-world image shift and achieving performance parity with ERM on mild-shift tasks.

Shuang Song, Muhammad Syafiq Mohd Pozi, Nik F. Farid · 0 citations
Jul 2026

Explainable Deepfake Detection Challenge

The Explainable Deepfake Detection Challenge at ACM Multimedia 2026 is designed to benchmark this joint capability of classification metrics with semantic similarity, simplicity, and intent-aware grounding metrics that assess whether explanations identify the relevant manipulated entities and supporting visual evidence...

Abhijeet Narang, Kartik Kuckreja, Shreya Ghosh et al. · 1 citation
Preprint Aug 2026

ConceptTS: LLM-Guided Concept Bottlenecks for Interpretable Multivariate Time-Series Forecasting

ConceptTS is introduced, an interpretable forecasting framework that organizes its predictions around named, human-readable concepts that achieves accuracy competitive with strong black-box baselines while producing semantically meaningful concept activations.

Yichen Jiang, Yueqiao Chen, Dong-Yu Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.