Concept-Bottleneck Models with Expert-Discovered Versus Automatically-Discovered Concepts: A Comparative Study on Fine-Grained Fungal Classification
Abstract
Interpretable-by-design architectures, such as Concept-Bottleneck Models (CBMs), are essential for trustworthy AI under regulations like the EU AI Act. While classical CBMs rely on costly expert labels, recent automated methods use vision–language models for concept discovery. However, whether automation preserves interpretability remains an open question. We systematically evaluate four automated methods (LF-CBMs, LaBo, PCBMs, and VLG-CBMs) against an expert-supervised CBM using fine-grained binary classification (Amanita muscaria vs. Boletus edulis) from the FungiTastic dataset. Although automated methods match expert accuracy, they systematically fail with regard to two key interpretability requirements: concept atomicity and instance-level grounding. Specifically, automated concepts conflate multiple morphological properties, yield scores ungrounded in visual evidence, and produce biologically implausible class-concept associations; a specific failure mode, cross-class contamination, is one that the expert-supervised vocabulary avoids by construction, though expert supervision remains subject to its own annotation and validation risks. We characterize four distinct failure modes across these architectures, concluding that automated concept discovery cannot yet substitute domain expert supervision in safety-critical tasks where explanation fidelity is a functional requirement alongside predictive accuracy.