Aug 2026· Discover Life· Vol 56· 0 citations· 53 references
TL;DR
A hybrid Neural Network-Transformer encoder trained on four complimentary datasets maintains competitive accuracy across all four datasets while producing more retrieval-relevant embeddings and much more consistent performance under degraded queries suggests that the model may serve as a promising foundation for future CBMIR research and potential clinical exploration.
Abstract
Clinicians can quickly find previous cases that are visually and semantically similar to a query image using content-based medical image retrieval (CBMIR); however, its practical implementation is hindered by multiple modalities, a lack of labeled data, and vulnerability to real-world artifacts like cropping and blurring. This research presents a hybrid Neural Network-Transformer encoder trained on four complimentary datasets: the COVID-19 Radiography Database, the Kvasir GI endoscopy dataset, and the ImageCLEFmed 2007 and 2009 radiology benchmarks. The model creates compact 128-dimensional embeddings that integrate particular local cues with more general anatomical contexts using a ResNet-50 convolutional backbone, a lightweight transformer head, and a gated fusion module. In order to improve robustness and cross-domain generalization, this study employs data augmentation that mimics real corruptions during training and evaluates retrieval under both clean and intentionally degraded queries. Using a unified methodology based on Precision@K, Recall@K, and mean average precision (mAP), the augmented hybrid model maintains competitive accuracy across all four datasets while producing more retrieval-relevant embeddings and much more consistent performance under degraded queries. This suggests that the model may serve as a promising foundation for future CBMIR research and potential clinical exploration.
Medical image analysis has undergone transformative progress with the application of deep learning models. However, existing architectures often struggle to effectively balance local feature extraction with global contextual understanding, which is crucial for complex diagnostic tasks such as Retinopathy of Prematurity (ROP) detection. In this study, we present a pretrained lightweight Conformer model tailored for medical image classification. The model integrates convolutional layers for capturing fine-grained spatial features with transformer blocks that capture long-range dependencies, creating a unified architecture capable of robust representation learning. We evaluate the model across multiple benchmark medical imaging datasets, including ROP, BloodMNIST, RetinalMNIST and other MedMNIST benchmark datasets. With 93.61% accuracy on the ROP dataset and 99.12% accuracy on BloodMNIST, experimental results show competitive classification performance while lowering model complexity to 12.4 million parameters and 3.2 GFLOPs. Experimental results demonstrate that the comparative studies versus CNN-based and transformer-based architectures, such as ResNet50, Swin-Tiny, ConvNeXt-Tiny, Vision Transformer, and MedViT. The findings show that in clinical settings with limited resources, the suggested lightweight Conformer offers a practical and computationally efficient alternative for medical image interpretation. Furthermore, the lightweight design ensures computational efficiency, making it suitable for deployment in resource-constrained healthcare environments. These findings validate the lightweight Conformer model’s potential for scalable, accurate, and real-time medical image classification.
Sreelekshmi Vijayasree, Adithya K. Krishna, Akarsh S. Nair et al.· Journal of Imaging· 0 citations
Medical image segmentation requires high accuracy and robustness, yet practical commercial deployment also demands privacy preservation and computational efficiency. In this context, the U-Net architecture, which can be inherently decoupled into independent encoder and decoder components, serves as a natural commercial choice. However, pure Transformer-based variants like Swin-UNet often suffer from insufficient local detail capture and limited interpretability. In this paper, we propose a lightweight hybrid architecture built upon the Swin-UNet framework. Our model integrates a parallel CNN encoder to complement the shallow layer reasoning of Swin Transformers with local texture features. To bridge the semantic gap and enhance fine-grained spatial detail recovery, we design an asymmetric feature fusion strategy and introduce cross-layer skip (XSkip) connections that explicitly propagate shallow CNN features into the decoder. We further incorporate novel loss functions and an auxiliary supervision head (Aux-Head) to strengthen training stability, boundary delineation, and intermediate feature interpretability. Extensive experiments on the Synapse multi-organ segmentation dataset demonstrate that our approach achieves state-of-the-art competitive Dice scores and Hausdorff distances, offering an accurate, efficient, and interpretable solution for clinical deployment.
TransCat, a hybrid CNN-Transformer architecture for medical image segmentation, is proposed and an extended deformable attention mechanism with attentive value identification is developed, to control the computational burden caused by the enlarged token set.
Jin Wang, Zheng-Hua Yang, Dong-Ming Zhou et al.· Frontiers in Bioinformatics· 0 citations
Medical imaging plays an essential role in the early detection and clinical assessment of multiple diseases; however, manual image interpretation is time-consuming and can be affected by inter-observer variability, particularly when subtle pathological patterns are present. Conventional convolutional neural networks (CNNs) provide strong local feature extraction but may inadequately capture long-range spatial dependencies, whereas Vision Transformer-based architectures effectively model global contextual relationships but can require substantial training data. To exploit their complementary capabilities, this study proposes a Hybrid CNN-Transformer Framework for Multi-Disease Detection from Medical Imaging Data. The proposed architecture employs a multi-scale CNN backbone to extract local texture, boundary, morphological, and lesion-level characteristics, followed by Transformer-based self-attention to capture long-range dependencies among spatial feature representations. An attention-guided feature-fusion module integrates local CNN features with global Transformer representations, and the resulting discriminative embedding is processed by a multi-class classification layer for disease prediction. Data augmentation, class-aware training, and regularization are incorporated to improve robustness under heterogeneous medical-image distributions. Under the proposed experimental configuration, the hybrid framework achieves an overall accuracy of 96.74%, sensitivity of 95.92%, specificity of 97.18%, precision of 96.31%, F1-score of 96.11%, and area under the receiver operating characteristic curve (AUC) of 0.986. Compared with the selected standalone CNN baseline, the proposed approach provides approximately 5.2% relative improvement in accuracy and 5.8% improvement in F1-score. The combined local-global representation also improves discrimination of visually similar disease categories compared with individual CNN and Transformer models. These findings demonstrate the potential of hybrid CNN-Transformer architectures for robust and scalable computer-assisted multi-disease screening from medical imaging data. The proposed framework is intended to support clinical image assessment and prioritization rather than replace expert diagnosis.
M. Balakrishnan, K. Ananthi, S. R. et al.· International journal of com...· 0 citations
Results indicate the WVM-UNet architecture effectively captures discriminative features for precise medical image segmentation, and demonstrates the competitive performance of the method on multiple public datasets.
Yulong Yang, Wenchao Gao, Zheng-Guo Wu et al.· Journal of Imaging· 0 citations
Accurate classification and efficient retrieval of histopathological images are essential for the diagnosis of lung adenocarcinoma (LUAD). Existing deep learning approaches for Content-Based Histopathological Image Retrieval (CBHIR) typically generate single-scale embeddings, missing the richer spatial context from earlier network stages. We propose a unified framework composed of: (1) a ConvNeXt V2 backbone with an integrated Convolutional Block Attention Module (CBAM) for multi-scale feature extraction, (2) an Adaptive Weighted Fusion Neck with learnable softmax-normalized weights, and (3) a novel Hybrid Representation Head producing an 18,496-dimensional descriptor by concatenating global, spatial, and attention-weighted features. Evaluated on the WSSS4LUAD dataset (10,087 patches, four tissue classes), our model achieves 85.03% accuracy (5-fold CV: 83.35 ± 0.93%), F1-score of 0.8116, mean Average Precision (MAP) of 0.8323 for retrieval, and an Expected Calibration Error (ECE) of 0.0378. Ablation experiments confirm that all proposed modules contribute positively, with the Attention Branch being the most impactful (Δ = −2.13%). The framework further provides Gradient-weighted Class Activation Mapping (Grad-CAM) explainability for clinical interpretability.
Noor Ahmed, F. Mahan, J. Karimpour· Computers· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.