Hybrid CNN-Transformer Based Brain Tumor Detection and Classification Using MRI: A Novel Framework with Multi-Level Attention and Regional Explainability
Aug 2026· International journal of computer information systems and industrial management applications· Vol 18, pp. 1115-1130· 0 citations
TL;DR
A novel hybrid convolutional neural networks and transformer architecture, HybCT-Net, augmented with a multi-level attention module and a regional explainability pipeline for brain tumor detection and classification is proposed, demonstrating superior performance than contemporary CNN, transformer and hybrid baselines.
Abstract
Automated and accurate classification of brain tumors from MRI (magnetic resonance imaging) for clinical applications is essential. However, it still poses a challenge owing to multiple factors. This includes the high intra-class heterogeneity, vague tumor boundaries, and lack of transparency in diagnostic reasoning. In this paper, we propose a novel hybrid convolutional neural networks and transformer architecture, HybCT-Net, augmented with a multi-level attention module and a regional explainability pipeline for brain tumor detection and classification. The framework employs local feature extraction of deep CNN encoders and the long-range dependency modeling capacity of lightweight vision transformers in conjunction. The MLAM incorporates attention mechanisms such as channel-wise squeeze-and-excitation gating, spatial convolutional block attention, and patch-based multi-head self-attention to enhance salient features at multiple semantic scales. A Hybrid Feature Fusion (HFF) is a block for adaptive fusion of the CNN feature maps and the transformer tokens that bridges the semantic gap between the convolution-based and attention-based features [24]. Moreover, the REP combines Gradient-weighted Class Activation Mapping with transformer attention rollout maps for producing regional heatmaps at pixel-wise spatial detail, enhancing clinical trust. The efficacy of the proposed model is established through extensive experiments on a multi-class brain MRI dataset with glioma, meningioma, pituitary tumor and no tumor. HybCT-Net attains a classification accuracy of 98.78%, with a macro-F1 score of 98.70% and an AUC of 0.9943. Comparative experiments demonstrate superior performance than contemporary CNN, transformer and hybrid baselines. The contributions of various architectural components are validated using ablation studies and computational analysis shows deployment feasibility. Qualitative visualizations suggest that the regional explainability maps correlate well with boundaries of the tumor as annotated by radiologists. Thus, the framework has the potential for use in clinical decision support in the real world.
The proposed XAIViT framework has strong potential as an Explainable Artificial Intelligence (XAI)-based clinical decision support system for MRI-based brain tumor analysis and Gradient-weighted Class Activation Mapping-based visual explanations demonstrated that the model consistently focused on anatomically relevant tumor regions, thereby improving transparency and trustworthiness.
I. H. A. Wahab, M. Jamil, Rosihan Rosihan· Engineering, Technology &...· 0 citations
Results show that ViT–BiLSTM's classification performance is superior to those of traditional deep learning methods: among all the tumor categories its accuracy is higher, its fine-tuning more perfect, as well as, its Recall rates greater.
Nagham Salim Mohammed, Omar S. Almolaa, A. S. Abdullah et al.· ITEGAM- Journal of Engineeri...· 0 citations
Medical imaging plays an essential role in the early detection and clinical assessment of multiple diseases; however, manual image interpretation is time-consuming and can be affected by inter-observer variability, particularly when subtle pathological patterns are present. Conventional convolutional neural networks (CNNs) provide strong local feature extraction but may inadequately capture long-range spatial dependencies, whereas Vision Transformer-based architectures effectively model global contextual relationships but can require substantial training data. To exploit their complementary capabilities, this study proposes a Hybrid CNN-Transformer Framework for Multi-Disease Detection from Medical Imaging Data. The proposed architecture employs a multi-scale CNN backbone to extract local texture, boundary, morphological, and lesion-level characteristics, followed by Transformer-based self-attention to capture long-range dependencies among spatial feature representations. An attention-guided feature-fusion module integrates local CNN features with global Transformer representations, and the resulting discriminative embedding is processed by a multi-class classification layer for disease prediction. Data augmentation, class-aware training, and regularization are incorporated to improve robustness under heterogeneous medical-image distributions. Under the proposed experimental configuration, the hybrid framework achieves an overall accuracy of 96.74%, sensitivity of 95.92%, specificity of 97.18%, precision of 96.31%, F1-score of 96.11%, and area under the receiver operating characteristic curve (AUC) of 0.986. Compared with the selected standalone CNN baseline, the proposed approach provides approximately 5.2% relative improvement in accuracy and 5.8% improvement in F1-score. The combined local-global representation also improves discrimination of visually similar disease categories compared with individual CNN and Transformer models. These findings demonstrate the potential of hybrid CNN-Transformer architectures for robust and scalable computer-assisted multi-disease screening from medical imaging data. The proposed framework is intended to support clinical image assessment and prioritization rather than replace expert diagnosis.
M. Balakrishnan, K. Ananthi, S. R. et al.· International journal of com...· 0 citations
Brain tumor classification from Magnetic Resonance Imaging (MRI) remains challenging because of tumor heterogeneity, the overlapping intensity distribution of tumors, and the limited availability of medical images. This study introduces a lightweight deep learning model by integrating a Hybrid Channel–Spatial Attention (HCSA) module into the EfficientNet-B0 architecture for multi-class brain tumor classification. The HCSA module processes channels and spatial features sequentially to enhance feature representation by focusing on tumor-relevant information and better localizing the tumor region while maintaining computational efficiency. The proposed framework was evaluated on a publicly available Kaggle MRI brain tumor dataset, comprising 5,600 MRI images across 4 classes – glioma, meningioma, pituitary tumor and no-tumor. Comparative experiments were conducted with ResNet50, DenseNet121, EfficientNet-B0, and attention-enhanced EfficientNet models with identical preprocessing, data augmentation and training conditions. The proposed HCSA-EfficientNet-B0 model achieved a classification accuracy of 98.57%, an F1-score of 0.99, and an ROC-AUC value of 0.999 on the internal hold-out test set. Ablation results indicate that the combination of channel attention and spatial attention consistently outperforms each of the two attention mechanisms in isolation. Moreover, the proposed framework effectively captures the tumor-relevant regions, as suggested by Grad-CAM visualizations, which provide qualitative evidence supporting the interpretability of the classification results. The proposed framework is effective for multi-class brain tumor classification using MRI images and is able to improve classification performance while achieving computational efficiency.
K. M, D. V., G. Sreenivasulu et al.· Journal of Trends in Compute...· 0 citations
Magnetic resonance imaging (MRI) brain tumor classification is an important factor in the early diagnosis and plans of treatment. Nevertheless, the current methods of deep learning lack robustness to spatial variations, lack attention to tumor-relevant features, and high computational complexity, which limits their use in the real world. To solve these issues, the paper provides a proposal of DA-CBAM EfficientNet which is a rotation-augmented and compressed transfer learning architecture to achieve precise and efficient brain tumor classification. The given model relies on an EfficientNet backbone, which applies the transfer learning to utilize the available pretrained representations and adjust the domain specificities of MRI. Data augmentation is done in a systematic manner by using rotation-based to provide orientation-invariance learning and better generalization. The Convolutional Block Attention Module (CBAM) and a Dynamic Attention-Weighted Enhancement Function (DAWEF) are used to gain a better understanding of informative feature channels and spatial locations related to tumors that can be enhanced by learning a discriminative feature. To achieve feasibility of deployment, pruning and quantization are done in a structured way to reduce the size of the model and inference latency by a fair margin without affecting the classification performance. Extensive experimental testing of an experimental dataset of a benchmark brain MRI shows that the proposed DA-CBAM EfficientNet will always be the best among the current state-of-the-art approaches in terms of accuracy, robustness, and computational efficiency. The findings validate the adequacy of the proposed framework to real time clinical decision support and resource limited healthcare settings.
Chinmayi Tamirisa, Saziya Tabbassum· International journal of com...· 0 citations
One of the most aggressive and lethal forms of neoplasia is brain tumor because they are typically diagnosed at an advanced stage, and there are inherent surgical complexities associated with the neuro-anatomic environment. To fill this critical diagnostic gap, the new methodology has been proposed that will be used to accurately and early segment tumor boundaries in 3D volumes of MRI. The framework is a Multi-Scale Self-Supervised Hybrid (MSSH) that was optimized in a unique way by hierarchical neural synthesis. Although recent developments have pursued 3D Self-Supervised Contrastive Learning to detect glioblastoma and Vision-Language Models to diagnose multimodally. The design explicitly targets the capture of the multifaceted topology of structural and reactive neuro-oncological features. A hybrid architecture, inspired by ensemble learning, is used to effectively combine the localized strength of Convolutional Neural Networks (CNNs) with the scale relational context of 3D Transformers. A self-trained Masked Voxel Reconstruction (MVR) protocol is used to attain a good balance between high-fidelity feature extraction and spatial performance. The method has been used to augment better precision to pathological limits by blending structural inductive biases and global attention mechanisms. Moreover, a multi-scale attention-based fusion framework enables the system to obtain context-related information at diverse resolutions, which is quite useful in classifying complex brain pathologies. The combination of these hybrid modules, combined with the resource efficient mixed-precision training, is what ensures that the system is not only explainable, but also fully deployable on commodity-grade clinical hardware. Quantitative validation is used to confirm the accuracy of the framework, with the Dice Similarity Coefficient (DSC) showing a value of 0.938 and Hausdorff Distance (HD95) at 2.84mm. Finally, the proposed system supports patient-specific planning of surgery and enhances the overall prognosis through the provision of a comprehensive volumetric analysis of the 3D tumor data.
K. A, Lohithadevi M, Kaviya Suresh· 2026 4th International Confe...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.