Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Xception-Based Brain Tumor Classification with GAN Augmentation and Multi-Technique Explainable AI

Automated brain tumor classification is essential for early detection and treatment planning. This study addresses the critical gap in Bangladeshi population-specific diagnostic systems by developing a comprehensive framework that integrates Generative Adversarial Network (GAN)-based augmentation with deep transfer learning and explainable artificial intelligence for MRI-based brain tumor classification. The proposed approach utilizes the PMRAM dataset comprising 3,505 T1-weighted MRI images from Bangladeshi patients across four classes (glioma, meningioma, pituitary tumor, and normal). A custom deep convolutional GAN generates 300 synthetic images per class, resulting in 40% dataset augmentation. The Xception architecture with ImageNet pre-training is employed using a two-stage training strategy consisting of feature extraction followed by fine-tuning of the final 100 layers. Stratified 5-fold cross-validation is conducted to compare performance against DenseNet169, MobileNetV2, and InceptionV3, while multi-technique explainable AI (GradCAM++, Score-CAM, SHAP, and LIME) provides clinical interpretability. Quantitative XAI evaluation via attribution map IoU and deletion faithfulness metrics confirms that highlighted regions are decision-critical and spatially concordant. Experimental results show that the proposed Xception-based model achieves a mean accuracy of 98.08% ± 0.42%, outperforming DenseNet169 (97.66%), MobileNetV2 (97.40%), and InceptionV3 (96.98%). The model attains a mean AUC of 0.997 with an Expected Calibration Error of 0.0135, and statistical validation confirms significance (p = 0.036). Ablation studies further demonstrate the contributions of GAN augmentation (+3.33%), fine tuning (+6.88%), and transfer learning (+21.81%). The proposed framework provides a comprehensive benchmark for Bangladeshi brain tumor classification with accuracy approaching clinical grade performance and interpretable predictions, highlighting the effectiveness of GAN-based transfer learning for population specific medical AI systems.

S. M. Tawhid, Yash Rohan, Shochi Akter et al. · 0 citations
#large language models Review Open access Sep 2026

Multimodal medical diagnosis: a mini review of LLM–vision fusion models in low-resource healthcare settings

Recent advances in large language models (LLMs) and vision transformers have enabled multimodal systems that integrate clinical text with medical imaging for diagnostic decision-making. While these systems show promising results on benchmark datasets in well-resourced research settings, their applicability in low-resource healthcare environments where diagnostic disparities are most severe remains limited and poorly understood. This mini review synthesizes key developments in LLM–vision fusion architectures from 2018 to 2026, with a focus on radiology-oriented visual question answering (VQA) and report generation systems viewed from a deployment perspective. Rather than comprehensively cataloguing multimodal medical AI, we synthesize the evolution of LLM–vision fusion architectures and discuss complementary deployment-enabling strategies, including parameter-efficient adaptation, post-training quantization, federated learning, and multilingual support, where they directly improve the feasibility of radiology AI in resource-constrained healthcare settings. Rather than focusing solely on performance benchmarks, we examine these approaches through a deployment-oriented lens, highlighting trade-offs between representational capacity, computational efficiency, interpretability, and memory footprint. We argue that current progress remains substantially shaped by model scaling and benchmark optimization, which often do not address the memory, connectivity, and annotation constraints of low-resource healthcare systems. While cross-modal transformer architectures provide strong representational alignment, their computational demands and reliance on large curated datasets limit real-world deployment. In contrast, emerging directions including parameter-efficient fine-tuning, post-training quantization, federated learning, and modular agent-based systems offer more tractable pathways toward clinical integration under hardware and data constraints. To bridge the gap between benchmark performance and clinical utility, we identify concrete challenges in data scarcity, multilingual coverage, and calibration, and propose a shift toward lightweight, interpretable, and hardware-aware multimodal AI. This perspective highlights the need to move beyond scaling-centric design toward models that can run on 4–8 GB VRAM, operate offline, and generalize across languages and imaging equipment.

Kahakashan Ashraf, Md.Hamid Hosen, N. Farah et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.