Skip to content
Preprint

SegBanana: Steering Unified Multimodal Models into Medical Segmenters

Sep 2026 · 0 citations · 54 references
Computer Science

TL;DR

SegBanana is proposed, to the authors' knowledge, the first agentic visual generation framework for training-free medical image segmentation and builds on a frozen UMM as the core generative model, augmented with Anatomy-Aware Knowledge Retrieval and Comparative Quality Critique to unlock its potential segmentation capability.

Abstract

Medical image segmentation remains challenging in practical deployment, as models often struggle to generalize beyond the distributions covered by their training data and high-quality pixel-level annotations are typically unavailable for adaptation. Inspired by the cross-task transferability of large language models, we investigate whether unified multimodal models (UMMs) can transfer their pretrained visual understanding, reasoning, and generation capabilities to medical image segmentation without task-specific post-training. By recasting segmentation as structured visual generation, we find that frontier UMMs (e.g., Nano Banana) already exhibit basic segmentation capabilities across diverse clinical scenarios, but still struggle with challenging tasks requiring specialized anatomical or domain-specific knowledge. We further show that these limitations can be effectively mitigated by incorporating visual anatomical knowledge from in-context exemplars, expanding candidate solutions through repeated sampling, and refining suboptimal predictions via targeted editing.Motivated by these observations, we propose SegBanana, to our knowledge, the first agentic visual generation framework for training-free medical image segmentation. SegBanana builds on a frozen UMM as the core generative model, augmented with Anatomy-Aware Knowledge Retrieval and Comparative Quality Critique to unlock its potential segmentation capability. A State-Aware Multimodal Controller maintains structured state and iteratively orchestrates these tools, repeatedly refining intermediate predictions toward higher-quality masks. Across eight medical segmentation datasets, SegBanana achieves an average mDice of 77.45%, outperforming representative generalist (SAM3 and SegGPT) and medical-specific (BiomedParse and MedSAM3) baselines by at least 14.93 points, while remaining robust to out-of-domain visual supports.

View source

Similar papers

Conference 2026

GeoFuse-SAM: A Multimodal Data Fusion Framework for Boundary-Aware Foundation Model Adaptation in Medical Image Segmentation

In medical image segmentation, accurately identifying anatomical structure boundaries is a core task for clinical computer-aided diagnosis. Although vision foundation models, represented by the Segment Anything Model (SAM), have demonstrated strong generalization potential, their deployment in fully automated clinical...

Qi-Yuan Wang · 0 citations
Preprint Aug 2026

Curia-MAE: Multi-Modal Multi-Anatomy MAE Pre-Training for 3D Medical Image Segmentation

Radiology foundation models learn transferable representations that can be adapted to new tasks by training only small layers on top of a frozen encoder. Dense prediction tasks such as 3D segmentation are, however, underrepresented in their evaluation, and, with the encoder kept frozen, pre-trained models still fall sh...

Th'eo Danielou, A. Saporta, L. Alberge et al. · 0 citations
Preprint Aug 2026

B-MIM: Biased Masked Image Modeling for Generalizable Segmentation of Fine-Grained Anatomical Structures

B-MIM is introduced, a modification of the iBOT objective that stochastically reduces global semantic alignment to prioritize local patch reconstruction and suggests that reducing global semantic pressure during pretraining enhances generalization to intricate anatomical structures.

S. González, Karen Sanchez, J. M. Saavedra et al. · 0 citations
Preprint Aug 2026

SymmAdapt: Symmetrical Flow Matching for Source-Free Domain Adaptation in Medical Image Segmentation

This work introduces a novel SFUDA framework built on Symmetrical Flow Matching, a unified generative model that segments an input image and synthesizes a source-like image from a mask within the same learned flow that outperforms SFUDA baselines and is competitive with conventional UDA methods.

T. Grossman, N. Cahan, H. Greenspan · 0 citations
Sep 2026

FPGL: A meta-learned implicit neural representation framework for medical image segmentation.

Modern state-of-the-art deep learning architectures for medical image segmentation rely strictly on feed-forward passes over dense pixel/voxel grids, which scale poorly to large signals. Implicit neural representations (INRs) offer a lightweight, continuous alternative to raw grids, but are traditionally signal-specifi...

Kushal Vyas, Daniel Kim, T. Netherton et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.