Skip to content
Open access

CrossMeta-ViT: Concomitant Cross-Attention Guided Multimodal Fusion for Skin Lesion Classification

Aug 2026 · Alfa Mühendislik ve Uygulamalı Bilimler Dergisi · 0 citations · 17 references

Abstract

Skin cancer is one of the most prevalent and deadly malignancies, necessitating early and precise diagnosis to improve patient survival rates. While recent advancements in deep learning have produced robust computer-aided diagnosis (CAD) systems, the majority of these models rely exclusively on visual features extracted from dermoscopic or clinical images. Conversely, dermatologists synthesize visual cues with patient clinical metadata (e.g., age, gender, continuous bleeding, and itchiness) to reach an accurate diagnosis. Prior research attempting multimodal fusion has largely depended on black-box concatenation-based late-fusion strategies or simple neural compression modules, obscuring clinical reasoning. To bridge this gap toward Explainable Artificial Intelligence (XAI), we propose CrossMeta-ViT, a deep vision-transformer metadata-fusion framework designed for transparent skin lesion classification. The core contribution of this architecture is a Cross-Attention Fusion module. Instead of sequentially concatenating features, CrossMeta-ViT utilizes encoded patient tabular metadata as a Query (Q) to dynamically attend to visual image patches acting as Keys (K) and Values (V). This mechanism ensures that structural image features are selected and weighted under the direct guidance of the patient's clinical history. We evaluated the framework on the smartphone-captured PAD-UFES-20 dataset for binary classification (Benign vs. Malignant). Using 5-fold cross-validation, CrossMeta-ViT achieved a macro F1-score of 0.8723 ± 0.0009, malignant recall of 0.9927, specificity of 0.7651, and an AUC of 0.955, while the held-out single-split evaluation yielded an accuracy of 0.90. These results support the model as a recall-oriented and interpretable multimodal aid for telemedicine pre-triage rather than as a universally dominant classifier.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.