Aug 2026· Computer Methods and Programs in Biomedicine· Vol 286, pp.
109607
· 0 citations· 32 references
Medicine
TL;DR
These findings suggest that interpretability quality can be quantitatively assessed and improved through multi-method and ensemble-based analysis and enhances explanation robustness and supports reliable and transparent clinical decision support for prostate magnetic resonance imaging.
Abstract
Background
AND
Objective
Explainable artificial intelligence is essential for clinical adoption of deep learning models in prostate magnetic resonance imaging. Although ensemble learning can improve robustness, its impact on explanation stability, spatial consistency, and clinical interpretability remains insufficiently quantified. This study aims to evaluate post-hoc interpretability methods across both single-model and ensemble configurations, and to examine whether ensemble-based explanations provide more reliable and clinically meaningful insights than single-model explanations. Critically, this work treats interpretability as a measurable property rather than a purely qualitative visualization.
Methods
Convolutional neural networks and bagging-based ensemble models (five VGG16-based classifiers trained on bootstrap samples with replacement, aggregated by soft averaging) were trained on the public PROSTATEx dataset using T2-weighted and apparent diffusion coefficient images. Visual explanations were generated using Gradient-weighted Class Activation Mapping (Grad-CAM) and saliency maps. Lesion localization was evaluated using centroid distance and Dice similarity coefficient with expert-annotated lesion masks. An agreement metric was introduced to quantify spatial consistency between attribution methods and its relationship with prediction reliability.
Results
The baseline classifier achieved an area under the curve of 0.84, with sensitivity of 0.81 and specificity of 0.86. Grad-CAM localized lesion centroids with higher precision on T2-weighted images (mean error 6.93 pixels) than apparent diffusion coefficient images (mean error 16.3 pixels). Combining saliency maps and Grad-CAM improved the mean Dice score from 0.42/0.45 (individual methods) to 0.52. Ensemble-based explanations were significantly smoother and less variable than individual classifier explanations (Mann-Whitney U, Levene test, all p<0.001). The agreement metric strongly separated correctly and incorrectly classified cases (Mann-Whitney U=84.5, p<0.001; point-biserial r=0.749; ROC-AUC =0.960).
Conclusions
These findings suggest that interpretability quality can be quantitatively assessed and improved through multi-method and ensemble-based analysis. The proposed agreement-driven framework enhances explanation robustness and supports reliable and transparent clinical decision support for prostate magnetic resonance imaging.
The results demonstrate that combining adaptive preprocessing, patient-wise evaluation, and explainable deep learning holds promise for MRI-based Parkinson’s disease detection under a preliminary, dataset-specific evaluation, though substantial performance variability remains across different subject selections, rather...
Ioana-Teodora Isar, Nirvana Popescu· Algorithms· 0 citations
A unique explainable deep learning model that integrates Tiny-ConvNeXt and DenseNet169 using multi-stage brain tumor classification is proposed, indicating that the suggested model provides a transparent and reliable framework for brain tumor detection, with potential applications in practical clinical decision support...
Md Sadi Al Huda, K. Tanvir, Zubaida Akhter et al.· Artificial Intelligence and...· 0 citations
Introduction Skin cancer is among the most prevalent and life-threatening malignancies worldwide. Early and accurate detection significantly improves therapeutic outcomes. Automated classification of dermoscopic skin lesions remains challenging due to class imbalance, inter-class visual similarity, and lack of interpre...
NE. Sravani, Srinivas Koppu· Frontiers in Public Health· 0 citations
Despite the impressive diagnostic accuracy achieved by deep convolutional neural networks (CNNs), their decisionmaking processes are often challenging to interpret, which hinders clinical trust and widespread adoption. This paper introduces an interpretability framework based on activation patching for classifying brai...
Dharanika S, Dharsana Dharani Vg, Ramala T et al.· 2026 International Conferenc...· 0 citations
Background Multiparametric magnetic resonance imaging (mpMRI) is an established component of prostate cancer diagnosis; however, PI-RADS interpretation remains partly subjective, particularly for equivocal lesions. Radiomics can quantify imaging characteristics that may not be readily appreciated visually, while explai...
Sacheena Kumbara, Sudha Y. Soudi, Pramod Kumar V· EPRA international journal o...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.