Skip to content
Preprint

Hierarchical Prompt Injector for Domain Generalization Segmentation

Sep 2026 · 0 citations · 66 references
Computer Science

TL;DR

Considering the difficulty of learning spatially and semantically aware prompt injection, the Hierarchical Prompt Injector is proposed, which enables spatially adaptive prompt injection in foundation models and auxiliary supervision to align hierarchical prompts with their corresponding object regions is introduced.

Abstract

Domain Generalized Semantic Segmentation (DGSS) is a challenging task, as vision models often rely on low-level appearance cues that change across domains. In contrast, structural attributes exhibit cross-domain stability, motivating the use of structural priors for DGSS. Existing methods use prompt learning to transfer such priors into DGSS models, but typically encode each class as a single holistic prompt. Moreover, these methods apply prompts uniformly to all pixels, offering no mechanism to adapt when only a subset of object regions is visible due to viewpoint changes, occlusion, and environmental variation. We address this with \textbf{Spatial Hierarchical Prompts (SHP)} that enrich each class with region-level geometric anchors capturing structural appearance from distinct viewing angles, ensuring complementary coverage under arbitrary viewpoints. Additionally, we propose the \textbf{Hierarchical Prompt Injector (HPI)}, which enables spatially adaptive prompt injection in foundation models. HPI spatially grounds prompts by modeling their semantic relevance and spatial influence with visual features. Considering the difficulty of learning spatially and semantically aware prompt injection, we further introduce auxiliary supervision to align hierarchical prompts with their corresponding object regions. We achieve 70.62\% and 72.74\% mIoU on synthetic-to-real and real-to-real benchmarks, respectively. Code and checkpoints are released at https://github.com/MosukFate/HPI

View source

Similar papers

Conference Open access Sep 2026

F³S: feature fused few-shot segmentation with CLIP guided semantic priors and Sinkhorn attention refinement

Few-shot segmentation (FSS) remains a significant challenge due to the scarcity of annotated data and the need for precise object localization across novel classes. Existing approaches often rely on single-backbone architectures and coarse priors, which struggle to capture detailed semantics and precise spatial alignme...

Guo-Hua Geng, Xiao-Feng Wang · 0 citations
Preprint Aug 2026

LASA: Language-and-Source-Anchored Alignment for Domain Generalized Semantic Segmentation

The Language-and-Source-Anchored Alignment (LASA) framework is proposed, which comprises three synergistic components: Text-and-Source-Guided Style Transfer (TSGST), Domain-Aware Query Adapter (DAQA), and Domain-Aware Decoder Optimizer (DADO).

Jin-Hong Zhu, Wei-Qi Yan, Sheng-Chuan Zhang et al. · 0 citations
Aug 2026

MSPD-net: structural–appearance prototype decoupling for weakly supervised semantic segmentation

A complementary prototype representation framework is proposed, employing three modules to collaboratively improve pseudo-label quality and improves the discriminative ability of confused categories by generating semantically similar sub-category negative samples.

Wei Cao, Yong Jiang, Ruiying Wang · 0 citations
Preprint Sep 2026

Diffuse2Seg: Diffusion Models Can Segment Anything Without Supervision

Open-world entity segmentation aims to predict masks for arbitrary objects across domains and at multiple granularities, from parts to whole objects. In this setting, SAM sets a strong standard: trained on SA-1B, comprising 11M images and over 1B carefully annotated masks, it achieves remarkable zero-shot performance....

Christoph Hümmer, Joachim Sicking, Fabian Hüger et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation

We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absenc...

J. del Pino, Salvador Rodríguez, Alejandro Garabito et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.