Skip to content

CERD3D-UNet: context-enhanced residual 3D U-Net with dual-attention gates and hyperparameter optimization for multimodal brain tumor segmentation

Aug 2026 · The Visual Computer · Vol 42 · 0 citations · 70 references

TL;DR

CERD3D-UNet is introduced, a context-enhanced residual-dense 3D U-Net model with dual-attention mechanisms and hyperparameter optimization for accurate delineation of whole tumor, tumor core (TC), and enhanced tumor (ET).

View source

Similar papers

Open access Jul 2026

A 3D residual U-Net with attention-driven spatial pyramid pooling for accurate multimodal MRI brain tumor segmentation

Medical imaging plays a crucial role in the accurate detection and localization of brain tumors, which is essential for effective clinical diagnosis and treatment planning. However, conventional segmentation approaches often struggle to capture complex spatial dependencies in volumetric data. To address this limitation, this study proposes an enhanced 3D U-Net architecture for multi-modal MRI-based brain tumor segmentation. The proposed model leverages three-dimensional convolutional operations to effectively capture contextual and spatial information from volumetric inputs. Additionally, an automated preprocessing pipeline, including image resizing, intensity normalization, and data augmentation, is incorporated to improve model robustness and generalization. The performance of the proposed model is evaluated against a conventional U-Net and a ResNet-based segmentation model using standard metrics such as Dice coefficient, accuracy, Intersection-over-Union (IoU), precision, recall, and F1-score. Experimental results demonstrate that the proposed 3D U-Net achieves superior performance, with a Dice coefficient of 0.83 and a Jaccard index of 0.82, outperforming baseline models across all evaluation metrics. Furthermore, the model exhibits improved convergence behavior and reduced overfitting, indicating strong generalization capability. These findings highlight the effectiveness of the proposed approach for volumetric medical image segmentation. Future work will focus on optimizing hyperparameters, enhancing architectural design, and validating the model on larger and more diverse clinical datasets.

Retinderdeep Singh, C. Prabha, Navita Gupta et al. · 0 citations
Open access Aug 2026

REC-UNet: a 2D U-Net model with residual cross-dimensional attention for liver tumor segmentation

A novel residual “Enhancement-Calibration” U-Net architecture, termed REC-UNet, which achieves high overall segmentation accuracy across diverse lesion sizes and contrast conditions without relying on explicit size-stratified optimization.

Zhiyuan Wang, Lijun Liang, Wei Wu et al. · 0 citations
Preprint Aug 2026

Text-Guided Refinement of Multi-sequence Glioma Subregion Segmentation with a Vision-Language Foundation Model

Background: Accurate glioma subregion delineation is important for radiotherapy planning and longitudinal monitoring, but manual contour correction is time-consuming. Models such as nnU-Net may generalize imperfectly and lack clinician-directed text correction. Purpose: We investigated adapting a three-dimensional (3D) vision-language foundation model for text-guided brain tumor segmentation refinement. Methods: We developed a lightweight VoxTell-based framework. Pretrained VoxTell generated initial masks. Oracle prompts derived from segmentation errors encoded target, action, location, imaging evidence, edit size, and preservation constraints. Frozen Qwen/VoxTell prompt embeddings were injected through trainable projections into its multiscale decoder conditioning; other weights remained frozen. Training, validation, and testing used 901, 100, and 250 BraTS-GLI cases. Cross-dataset transfer was evaluated on 100 meningioma, metastasis, pediatric tumor, and UPENN-GBM cases. Results: On the internal test set using post-contrast T1-weighted input, correct instructions improved subregion Dice similarity coefficient (DSC; enhancing tumor, edema, and necrotic/non-enhancing core) from $0.774\pm0.158$ to $0.796\pm0.137$. They outperformed blank prompts ($0.762\pm0.155$; Holm-adjusted $p<0.001$, $d_z=0.71$) and contradictory prompts ($0.770\pm0.163$; $p<0.001$, $d_z=0.48$). In cross-dataset testing, correct instructions improved DSC from $0.527\pm0.287$ to $0.550\pm0.278$ and outperformed contradictory instructions ($0.504\pm0.275$; $p<0.001$, $d_z=0.43$). Conclusion: A 3D vision-language foundation model can perform instruction-guided refinement of glioma subregion segmentations. Sensitivity to correct, blank, and contradictory prompts suggests text-dependent contour editing rather than nonspecific post-processing, supporting further evaluation as a clinician-in-the-loop tool.

Zach Eidex, Yunyan Lin, M. Safari et al. · 0 citations
Open access Aug 2026

Multi-plane attention guided 3D nnU-Net for MRI-based cervical cancer segmentation

Objective Over 300,000 people die from cervical cancer, which is the fourth most frequent malignancy in women worldwide. Early identification of cervical cancer has been related to much greater survival rates, and the illness is usually preventable. Accurate cervical segmentation using magnetic resonance imaging (MRI) is critical for diagnosis, treatment planning and response assessment especially in image-guided brachytherapy. Methods In terms of 3D MRI cervical tumor image segmentation task, several cutting-edge state-of-the-art architectures are explored. We propose a novel model having MultiPlane Attention Guided Enhanced nn-UNet3D model that uses volumetric contextual information to learn spatial relationships across slices simultaneously improving border localization and noise robustness. Results It is applied on publicly available TCGA-CESC dataset which is part of the Cancer Genome Atlas program (TCGA) and it focuses on cervical squamous cell carcinoma and endocervical adenocarcinoma (CESC). Preliminary results show the accuracy, IoU and Dice score values of 95.04%, 87.87% and 93.52% respectively. Conclusion This study discusses the significance of MRI scans, which provide high-resolution anatomical information, improving the accuracy of segmentation algorithms in identifying and characterizing aberrant cervical tissues. Future research will concentrate on domain generality across scanners and institutions, multimodal MRI integration and clinical validation in prospective therapy scenarios.

Jheelam Mondal, Rajdeep Chatterjee, M. Gourisaria et al. · 0 citations
Open access Aug 2026

Bridging global context and local precision using a disagreement-based region specific ensemble of Swin UNETR and SegResNet for 3D glioma segmentation

Brain tumor is one of the most challenging neurological diseases to diagnose and even a minor inaccuracy in the tumor characterization can be fatal. An accurate and reliable brain tumor segmentation from 3D MRI images is a fundamental requirement for an effective diagnosis, treatment planning and assessment of outcome in neuro-oncology. Due to infiltrative growth of tumors, heterogeneity in its structure and diffuse boundaries of tumor regions, brain tumor segmentation is quite critical and challenging. Even a minor error in delineation can adversely affect surgical resection and radiotherapy planning. To address these challenges, this study proposes a region-adaptive ensemble framework that integrates the complementary strengths of two capable 3D segmentation models, SegResNet and Swin UNETR through a staged fusion strategy: simple averaging, region-adaptive soft weighting (RSW), and a disagreement-based region-specific refinement (DRE) for high-conflict voxels. The CNN-based SegResNet is capable in capturing fine-grained local textures and well-defined tumor cores due to its convolutional local bias whereas Transformer-based Swin UNETR is capable in modeling long range contextual dependencies across MRI volume due to its hierarchical Transformer architecture. These two models are finetuned on BraTS 2020 dataset and then integrated using a dynamic voxel-wise disagreement-based fusion strategy that adaptively weights model predictions based on parameters like regional confidence, historical performance and level of disagreement. The multi-run experimental evaluation of the architecture on BraTS 2020 dataset is able to achieve impressive dice scores of 0.9447 ± 0.0021 in Whole Tumor (WT), 0.9231 ± 0.0033 in Tumor Core (TC) and 0.9071 ± 0.0041 in Enhancing Tumor (ET) regions. These results indicate that the Region Adaptive fusion with a disagreement-based refinement between convolutional and transformer-based models leads to a robust framework for brain tumor segmentation, that upon further research and validation might turn out to be suitable for clinical decision making and treatment planning.

Amrit Baskota, Shubham Ghimire, Baskaran P · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.