Introduction The automated classification of dermoscopic skin lesions is inherently challenging due to pronounced class imbalance, minimal inter-class variance, visual similarity across lesion types, and the requirement for clinically interpretable predictive outcomes. Methods The present study designed a heterogeneous Vision Transformer ensemble framework for seven-class skin lesion classification using the HAM10000 dataset. The framework integrates three architecturally distinct backbones Swin-Tiny, ViT-Base, and DeiT-Small enhanced with a novel Regional Attention Wrapper (RAW) for spatially selective feature aggregation. The generated outputs are combined via a stacking protocol wherein a trained MLP meta-learner resolves class-aware disagreements among the models. Class imbalance is addressed using class-adaptive augmentation with class-weighted focal loss and MixUp regularisation. Results The proposed framework achieved 98.37% accuracy, weighted F1-score of 98.39%, and mean AUC of 0.999, surpassing all three individual backbones across all metrics. MEL misclassifications were reduced by 78% compared to the weakest baseline, confirmed by McNemar's test (p < 0.0001). Comprehensive ablation studies validate the contributions of the RAW module, each ensemble component, and each augmentation strategy. Three-fold cross-validation yields a mean accuracy of 96.14 ± 0.36% and bootstrap confidence interval analysis confirms the reliability and reproducibility of the reported results. Discussion An exhaustive explainability framework comprising Regional Attention Maps, GradCAM++, SHAP, and t-SNE provides complementary spatial, gradient-based, pixel-level, and embedding-level interpretability, ensuring clinical trust, transparency, and trustworthiness expected from an automated dermoscopy system.
N. Sravani, Srinivas Koppu· Frontiers in Medicine· 0 citations
Kidney abnormalities, including cysts, tumors, and stones, are the most common renal disorders that can lead to severe complications such as chronic kidney disease or renal failure. Deep learning-based medical image analysis offers an effective approach for the accurate classification of kidney abnormalities, aiding the early diagnosis of renal disorders. However, its centralized training leads to inadequate privacy protection.
Considering the importance of ensuring individuals' data privacy, this study proposes a novel federated transfer learning framework for accurate classification of renal abnormalities using 12,446 kidney CT scan images and simultaneously preserves data privacy. CT scan images were preprocessed by resizing and normalization, followed by data augmentation techniques, including random rotations (±30°), horizontal flips, and color jitter, to address class imbalance and improve model generalization. Five pre-trained deep learning models such as MobileNetV2, EfficientNetV2-S, ResNet50, DenseNet121, and InceptionResNetV2 were trained across seven federated clients. Federated weighted averaging was employed for aggregation, and AES-256 encryption in CBC mode was applied to all model parameter transmissions between clients and the server.
MobileNetV2 achieved the best performance, attaining 99.48% accuracy, 99.29% precision, 99.32% recall, 99.3% F1-score, 0.9999 AUC-ROC, and log loss of 0.0247. Cross-client validation produced an average accuracy of 98.85% with a generalization gap of only −0.0063, indicating strong generalization across client datasets.
The proposed framework provides an effective balance between privacy preservation and communication efficiency, highlighting its potential for deployment in distributed clinical environments for kidney disease diagnosis.
Sai Sri Hemantha Konala, Srinivas Koppu· Frontiers in Artificial Inte...· 0 citations
Introduction Kidney-related disorders are one of the global health concerns that require timely detection to prevent severe health complications. The use of computed tomography (CT) images for accurate classification of kidney diseases is important. However, it is challenging to differentiate between classes due to the subtle visual differences. This study introduces a novel two-stage deep learning architecture that integrates self-supervised representation learning with supervised classification for kidney CT image analysis using a publicly available kidney CT image dataset. Methods In the first stage, the DINO framework with a Data-efficient Image Transformer (DeiT-Tiny) backbone is used to learn useful features from kidney CT images independent of labels. In the second stage, the pre-trained model is fine-tuned using labeled data to classify kidney abnormalities. To ensure model transparency and clinical trustworthiness, two explainable AI techniques are applied. Grad-CAM++ is used to highlight important regions contributing to predictions in kidney CT images. In addition, DINO’s inherent multi-head self-attention mechanism is analyzed across all attention heads to capture diverse attention patterns. Results and discussion Experimental findings indicate that the proposed framework achieves strong classification performance, with a test accuracy of 99.16%, AUC-ROC of 99.99%, F1 score of 98.97%, precision of 98.90%, and recall of 99.05%, while also providing clear interpretability for automated kidney disease classification. External validation on a CT dataset from Iraq has yielded 97.03% accuracy, supporting the generalizability of the proposed framework.
Sai Sri Hemantha Konala, Srinivas Koppu· Frontiers in Medicine· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.