Skip to content

HybridDermNet: A CNN–Transformer Framework with Latent Bottleneck Learning for Robust Skin Cancer Classification

Sep 2026 · Natural Resources for Human Health · 0 citations · 37 references

TL;DR

A CNN–Transformer-based framework with latent bottleneck learning for robust multi-class skin cancer classification is proposed, which offers compact, interpretable and generalizable representations for reliable dermatology decision support in heterogeneous imaging conditions.

Abstract

The classification of skin lesions from dermoscopic and clinical images remains challenging due to visually similar lesion categories, variability in image acquisition, class imbalance, and limited cross-dataset generalisation, which can undermine the reliability of diagnostic results. Current CNN-based approaches are good at local texture, pigmentation, and border information, while transformer models are good at global structure and long-range dependencies. Most hybrid methods, however, use simple concatenation or additive fusion, or rely on multi-backbone architectures that are very computation-intensive, produce redundant features, and risk overfitting. In this work, a CNN–Transformer-based framework with latent bottleneck learning for robust multi-class skin cancer classification is proposed. The algorithm locally and globally extracts features in parallel, calculates bidirectional attention between features, adaptively fuses features, encodes features in a compact latent space, optionally reconstructs features, and predicts with confidence. A composite objective combines classification, reconstruction, cross-feature consistency, latent regularisation, and transformation-consistency loss. HybridDermNet achieved 96.18% accuracy, 94.86% macro-F1, and 97.92% AUROC on HAM10000, and 95.42% accuracy, 93.92% macro-F1, and 97.46% AUROC on ISIC 2018. When transferring between datasets, it achieved 87.24% and 87.61% macro-F1, respectively, and, in external validation on PAD-UFES-20, achieved 80.67% macro-F1. The impact of attention-guided fusion and latent compression was verified through ablation, calibration, explainability, efficiency, and statistical analysis. The framework offers compact, interpretable and generalizable representations for reliable dermatology decision support in heterogeneous imaging conditions.

View source

Similar papers

Open access Sep 2026

Hybrid CNN–BiLSTM with Multi-Transformer Stacking for Skin Lesion Classification

Skin cancer is one of the most common diseases worldwide and, if left untreated, it can be life threatening. In this work, we propose a dual-branch deep learning framework that integrates a CNN–BiLSTM module for local spatial–sequential feature modeling with multiple transformer models (ViT, DeiT, SwinV2, and BEiT) for...

Maryem Zahid, Mohammed Rziza, Rachid Alaoui · 0 citations
Open access Sep 2026

MedFuse: dual-stream fusion of convolutional and vision transformer-based features for enhanced medical image classification

Medical image classification is fundamental to computer-aided diagnosis. Limited labeled samples, subtle inter-class differences, and heterogeneous lesion morphology make it difficult for a single representation to capture all relevant cues. This study investigates MedFuse, a simple dual-stream framework that combi...

Ya-Jing Ren, Hai Ling, Zheng Gu et al. · 0 citations
Open access Aug 2026

Hybrid CNN-Transformer Framework for Multi-Disease Detection from Medical Imaging Data

The proposed framework is intended to support clinical image assessment and prioritization rather than replace expert diagnosis, and demonstrates the potential of hybrid CNN-Transformer architectures for robust and scalable computer-assisted multi-disease screening from medical imaging data.

M. Balakrishnan, K. Ananthi, S. R. et al. · 0 citations
Aug 2026

Hybrid CNN–Transformer Architecture for Multi-Organ Disease Classification

Convolutional neural networks (CNNs) and Vision Transformers (ViTs) offer complementary strengths for medical image analysis: CNNs excel at capturing local texture and edge information through their inductive spatial bias, while transformers capture long-range dependencies through global self-attention but typically re...

S. Thanekar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.