SegBanana is proposed, to the authors' knowledge, the first agentic visual generation framework for training-free medical image segmentation and builds on a frozen UMM as the core generative model, augmented with Anatomy-Aware Knowledge Retrieval and Comparative Quality Critique to unlock its potential segmentation capability.
Abstract
Medical image segmentation remains challenging in practical deployment, as models often struggle to generalize beyond the distributions covered by their training data and high-quality pixel-level annotations are typically unavailable for adaptation. Inspired by the cross-task transferability of large language models, we investigate whether unified multimodal models (UMMs) can transfer their pretrained visual understanding, reasoning, and generation capabilities to medical image segmentation without task-specific post-training. By recasting segmentation as structured visual generation, we find that frontier UMMs (e.g., Nano Banana) already exhibit basic segmentation capabilities across diverse clinical scenarios, but still struggle with challenging tasks requiring specialized anatomical or domain-specific knowledge. We further show that these limitations can be effectively mitigated by incorporating visual anatomical knowledge from in-context exemplars, expanding candidate solutions through repeated sampling, and refining suboptimal predictions via targeted editing.Motivated by these observations, we propose SegBanana, to our knowledge, the first agentic visual generation framework for training-free medical image segmentation. SegBanana builds on a frozen UMM as the core generative model, augmented with Anatomy-Aware Knowledge Retrieval and Comparative Quality Critique to unlock its potential segmentation capability. A State-Aware Multimodal Controller maintains structured state and iteratively orchestrates these tools, repeatedly refining intermediate predictions toward higher-quality masks. Across eight medical segmentation datasets, SegBanana achieves an average mDice of 77.45%, outperforming representative generalist (SAM3 and SegGPT) and medical-specific (BiomedParse and MedSAM3) baselines by at least 14.93 points, while remaining robust to out-of-domain visual supports.
This paper proposes ReG-SAM, a SAM-based framework tailored to 2D vessel segmentation that leverages reference graph set for enhancing vascular representations and introduces two modality-aware representations derived from the reference masks.
Donghang Lyu, Zi-Chen Zhang, O. Dzyubachyk et al.· 0 citations
In medical image segmentation, accurately identifying anatomical structure boundaries is a core task for clinical computer-aided diagnosis. Although vision foundation models, represented by the Segment Anything Model (SAM), have demonstrated strong generalization potential, their deployment in fully automated clinical...
Qi-Yuan Wang· Poster Volume 0007 The 2026...· 0 citations
Radiology foundation models learn transferable representations that can be adapted to new tasks by training only small layers on top of a frozen encoder. Dense prediction tasks such as 3D segmentation are, however, underrepresented in their evaluation, and, with the encoder kept frozen, pre-trained models still fall sh...
Th'eo Danielou, A. Saporta, L. Alberge et al.· 0 citations
B-MIM is introduced, a modification of the iBOT objective that stochastically reduces global semantic alignment to prioritize local patch reconstruction and suggests that reducing global semantic pressure during pretraining enhances generalization to intricate anatomical structures.
S. González, Karen Sanchez, J. M. Saavedra et al.· 0 citations
This work introduces a novel SFUDA framework built on Symmetrical Flow Matching, a unified generative model that segments an input image and synthesizes a source-like image from a mask within the same learned flow that outperforms SFUDA baselines and is competitive with conventional UDA methods.
Modern state-of-the-art deep learning architectures for medical image segmentation rely strictly on feed-forward passes over dense pixel/voxel grids, which scale poorly to large signals. Implicit neural representations (INRs) offer a lightweight, continuous alternative to raw grids, but are traditionally signal-specifi...
Kushal Vyas, Daniel Kim, T. Netherton et al.· Medical Image Analysis· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.