Skip to content
Conference

Interpretable Brain Tumor MRI Classification Using Layer-Wise Activation Patching and Indirect Effect Estimation

Aug 2026 · 2026 International Conference on Modern Sustainable Systems (CMSS) · pp. 1367-1373 · 0 citations · 24 references

Abstract

Despite the impressive diagnostic accuracy achieved by deep convolutional neural networks (CNNs), their decisionmaking processes are often challenging to interpret, which hinders clinical trust and widespread adoption. This paper introduces an interpretability framework based on activation patching for classifying brain tumor MRIs, which identifies critical feature channels that influence decisions within a trained CNN. The indirect impact of each channel is assessed by observing the changes in classifier logits after substituting the activation of a chosen channel between a clean MRI slice and a contrasting slice at the third rectified convolutional stage. The resulting indirect-effect scores for each channel are utilized to re-weight globally pooled feature representations, which are then classified using a ridge-regularized logistic regression model. To enhance decision transparency while maintaining predictive performance, a convex combination of the CNN classifier and the logistic model is employed. The proposed framework is tested on two publicly accessible benchmark datasets: Br35H for binary tumor detection and the Nackara dataset for four-class brain tumor classification. Experimental outcomes from a deterministic training run yielded classification accuracies of 97.0% and 94.6% for the binary and four-class datasets, respectively The proposed framework provides an interpretable and auditable approach for brain tumor MRI classification. However, further validation using patient-independent, multi-center datasets acquired from different MRI scanners is required before clinical deployment.”,

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.