Aug 2026· Journal of imaging informatics in medicine· 0 citations· 14 references
Medicine
TL;DR
The results show that larger models and larger pretraining datasets do not automatically lead to better downstream performance, and transfer effectiveness in medical imaging is driven primarily by architectural inductive biases, pretraining strategy, and domain relevance.
TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.
Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al.· Frontiers in Bioinformatics· 0 citations
A comparative study of different deep learning architectures, including classical CNNs, deep hierarchical models, residual and dense networks, and compound-scaled architectures is presented, showing that deeper networks provide better representation, while residual connections and compound scaling improve training stability and efficiency.
Riyaz Mohammed· International Journal of App...· 0 citations
Medical imaging plays an essential role in the early detection and clinical assessment of multiple diseases; however, manual image interpretation is time-consuming and can be affected by inter-observer variability, particularly when subtle pathological patterns are present. Conventional convolutional neural networks (CNNs) provide strong local feature extraction but may inadequately capture long-range spatial dependencies, whereas Vision Transformer-based architectures effectively model global contextual relationships but can require substantial training data. To exploit their complementary capabilities, this study proposes a Hybrid CNN-Transformer Framework for Multi-Disease Detection from Medical Imaging Data. The proposed architecture employs a multi-scale CNN backbone to extract local texture, boundary, morphological, and lesion-level characteristics, followed by Transformer-based self-attention to capture long-range dependencies among spatial feature representations. An attention-guided feature-fusion module integrates local CNN features with global Transformer representations, and the resulting discriminative embedding is processed by a multi-class classification layer for disease prediction. Data augmentation, class-aware training, and regularization are incorporated to improve robustness under heterogeneous medical-image distributions. Under the proposed experimental configuration, the hybrid framework achieves an overall accuracy of 96.74%, sensitivity of 95.92%, specificity of 97.18%, precision of 96.31%, F1-score of 96.11%, and area under the receiver operating characteristic curve (AUC) of 0.986. Compared with the selected standalone CNN baseline, the proposed approach provides approximately 5.2% relative improvement in accuracy and 5.8% improvement in F1-score. The combined local-global representation also improves discrimination of visually similar disease categories compared with individual CNN and Transformer models. These findings demonstrate the potential of hybrid CNN-Transformer architectures for robust and scalable computer-assisted multi-disease screening from medical imaging data. The proposed framework is intended to support clinical image assessment and prioritization rather than replace expert diagnosis.
M. Balakrishnan, K. Ananthi, S. R. et al.· International journal of com...· 0 citations
Background/Objectives: Diabetic retinopathy is a major cause of preventable vision loss worldwide, making early and accurate disease grading crucial for timely treatment. Although convolutional neural network (CNN)- and transformer-based architectures have demonstrated promising performance for retinal image analysis, comprehensive comparisons within a unified experimental framework remain limited. This study systematically compares representative standard and lightweight CNN- and transformer-based architectures for multi-class DR grading. Methods: Six ImageNet-pretrained deep-learning models, including ResNet50, EfficientNet-B0, MobileNetV2, Vision Transformer (ViT), Swin-Tiny, and Swin Transformer, were evaluated on the APTOS 2019 retinal fundus image dataset under a unified experimental configuration with consistent preprocessing, data augmentation, training, and evaluation settings. All models were fine-tuned and evaluated independently over five runs with different random seeds. Their performance was assessed using accuracy, precision, recall, F1-score, area under the receiver operating characteristic curve (AUC), Quadratic Weighted Kappa (QWK), per-class analysis, computational efficiency, and statistical analysis. Results: Transformer-based models generally achieved higher mean classification performance than the evaluated CNN-based models. Swin-Tiny achieved the highest mean accuracy (82.3%), macro F1-score (64.4%), weighted F1-score (82.1%), and QWK (89.8%) across the five runs. Among the CNN-based models, EfficientNet-B0 achieved the strongest overall classification performance, whereas MobileNetV2 provided the lowest computational complexity. The results also highlighted differences in learning behavior and computational requirements across the evaluated architectures. Repeated experiments demonstrated stable performance across different random seeds, supporting the reliability of the proposed evaluation. Conclusions: Overall, this study provides a comprehensive comparison of representative CNN- and transformer-based architectures under consistent experimental settings and offers practical guidance for selecting suitable deep learning models for automated diabetic retinopathy screening.
Breast cancer has been one of the major causes of cancer mortality in the world. According to WHO report 2.5 million deaths are predicted as a result of breast cancer in the world in 2040. Although deep learning has demonstrated encouraging histopathology image analysis, the current methods frequently fail to provide local morphological information and global contextual information at the same time. In this paper, we present our HrybridViT-CAM, a hybrid deep learning system that integrates convolution neural networks( CNN ), Vision Transformers( ViT ) and multi-scale attention system in order to classify breast cancer using histopathology images. Its architecture has a two-way structure: a CNN arm (EfficientNetB7) to extract local features and a vision transformer arm to analyze the global context, fused together with a cross-attention fusion block. We use convolution block attention modules(CBAM) and deformal attention to improve feature discrimination. The model was tested on the BreaKHis dataset with various magnifications (40x, 100x, 200x, 400x) in binary as well as in multi-class classification. For binary classification (benign vs. malignant), HybridViT-CAM achieved accuracies of 99.87%, 99.42%, 98.95%, and 98.31% at 40χ, 100χ, 200x, and 400x magnifications, respectively. For eight-class subtype classification, the corresponding accuracies were 98.76%, 97.89%, 97.23%, and 96.54%, respectively. Grad-CAM++ and attention visualization techniques allowed explaining the results, which showed high correspondence(93.7% agreement) with pathologist diagnostic criterial. The proposed model was able to detect the malignant regions of interest (ROIs) like nuclei pleomorphism, atypia chromatin patterns and architectural distortions which are comparable to clinical diagnostic standards.
S. Angayarkanni, Mithila R, Koushik Rithik et al.· ITM Web of Conferences· 0 citations
Medical image analysis has undergone transformative progress with the application of deep learning models. However, existing architectures often struggle to effectively balance local feature extraction with global contextual understanding, which is crucial for complex diagnostic tasks such as Retinopathy of Prematurity (ROP) detection. In this study, we present a pretrained lightweight Conformer model tailored for medical image classification. The model integrates convolutional layers for capturing fine-grained spatial features with transformer blocks that capture long-range dependencies, creating a unified architecture capable of robust representation learning. We evaluate the model across multiple benchmark medical imaging datasets, including ROP, BloodMNIST, RetinalMNIST and other MedMNIST benchmark datasets. With 93.61% accuracy on the ROP dataset and 99.12% accuracy on BloodMNIST, experimental results show competitive classification performance while lowering model complexity to 12.4 million parameters and 3.2 GFLOPs. Experimental results demonstrate that the comparative studies versus CNN-based and transformer-based architectures, such as ResNet50, Swin-Tiny, ConvNeXt-Tiny, Vision Transformer, and MedViT. The findings show that in clinical settings with limited resources, the suggested lightweight Conformer offers a practical and computationally efficient alternative for medical image interpretation. Furthermore, the lightweight design ensures computational efficiency, making it suitable for deployment in resource-constrained healthcare environments. These findings validate the lightweight Conformer model’s potential for scalable, accurate, and real-time medical image classification.
Sreelekshmi Vijayasree, Adithya K. Krishna, Akarsh S. Nair et al.· Journal of Imaging· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.