Research on Dermoscopic Image Classification Method Integrating Multi-Scale Features and Attention Mechanism
To address the issues of lesion scale differences, easy loss of fine-grained pigment information, and background interference in dermoscopic images, this paper constructs a multi-scale attention fusion classification network MSAF-EfficientNetV2 based on EfficientNetV2-S. The network extracts features in four stages. AMSF achieves dynamic fusion through channel mapping, spatial alignment, and sample-related scale weights. LAA strengthens the main body of the lesion and irregular boundaries with channel-spatial attention, and uses weighted cross-entropy to alleviate the long-tailed distribution of the seven classes. The experimental results used publicly available data to form two levels of evidence: In the MedMNIST+ 224×224 end-to-end benchmark, the accuracy of DenseNet121 was 84.74±0.51%, and the AUC of DINO ViT-B/16 was 96.50±0.51%; in the histopathologically confirmed true melanoma ISIC_0000013, the Otsu dark region accounted for 22.29%, and the darkest cluster in Lab-K-means accounted for 16.60%, with a Dice concordance of 85.38%. The publicly available ISIC 2017 segmentation experiment further showed that the Dice/Jaccard ratio of VGG16-U-Net was 91.5%/84.6%, and the Dice/Jaccard ratio in this case reached 96.2%/92.6%. The total number of network parameters is 20.069 M and the number of FLOPs is 5.775 G, indicating that the newly added fusion and attention structures maintain controllable computational overhead.