ColoNet: DnCNN-EfficientNetV2-MLP Framework with Explainable AI for Colorectal Histological Texture Classification
Abstract
Abstract: Colorectal histological texture classification requires models capable of distinguishing visually similar tissue patterns while maintaining reliable predictive performance. We developed ColoNet as an integrated classification framework that combines Denoising Convolutional Neural Network (DnCNN)-based preprocessing, partial adaptation of an ImageNet-pretrained EfficientNetV2-RW-M backbone, and a compact Multilayer Perceptron (MLP) classifier with qualitative explainability. The Kather 2016 benchmark was used, comprising 5,000 hematoxylin and eosin (H&E)-stained tiles equally distributed across eight tissue classes. DnCNN was trained exclusively on the training data using synthetic noise, after which the processed images were analyzed by EfficientNetV2-RW-M with the last 15% of parameter tensors fine-tuned jointly with an MLP that maps 2,152-dimensional embeddings to eight class probabilities. Following exact and near-duplicate screening, a stratified image-level partition assigned 4,000, 500, and 500 tiles to the training, validation, and test sets, respectively. ColoNet achieved 97.20% test accuracy, with macro precision of 97.24%, recall of 97.22%, F1-score of 97.21%, specificity of 99.60%, Matthews correlation coefficient (MCC) of 96.82%, and area under the curve (AUC) of 99.38%. Bootstrap analysis provided a 95% confidence interval of 94.8 to 98.7% for accuracy. Compared with three classical classifiers and three end-to-end deep-learning baselines on the same test set, ColoNet achieved higher accuracy by 3.2 to 8.2 and 2.4 to 4.4 percentage points, respectively. Ablation analysis identified partial backbone fine-tuning and the nonlinear MLP head as the main contributors to performance, while denoising alone did not improve accuracy with a frozen backbone. Calibration, decision curve analysis, and complementary tile-level explanations using Gradient-weighted Class Activation Mapping (Grad-CAM), Grad-CAM++, Extended Gradient-weighted Class Activation Mapping (XGrad-CAM), Local Interpretable Model-agnostic Explanations (LIME), and SHapley Additive exPlanations (SHAP) further supported model assessment and practical interpretation of colorectal tissue patterns.