The Language-and-Source-Anchored Alignment (LASA) framework is proposed, which comprises three synergistic components: Text-and-Source-Guided Style Transfer (TSGST), Domain-Aware Query Adapter (DAQA), and Domain-Aware Decoder Optimizer (DADO).
Abstract
Domain Generalization Semantic Segmentation (DGSS) focuses on generalizing knowledge from labeled source domains to unseen target domains where data is unavailable during the training phase. While conventional methods utilize style randomization or feature normalization to mitigate domain shifts, they often impair feature integrity. Specifically, style randomization distorts the underlying feature manifold due to its coarse-grained nature, while feature normalization suppresses discriminative, domain-sensitive semantic details owing to its rigid design. To address these limitations, we propose the Language-and-Source-Anchored Alignment (LASA) framework, which comprises three synergistic components: Text-and-Source-Guided Style Transfer (TSGST), Domain-Aware Query Adapter (DAQA), and Domain-Aware Decoder Optimizer (DADO). Concretely, the TSGST module addresses manifold distortion by utilizing source features as structural anchors and vision-language model (VLM) priors as fine-grained guidance. To restore suppressed discriminative and domain-sensitive details, the DAQA module recalibrates object queries via categorical guidance and domain-aware signatures, while the DADO module aligns the resulting query distributions with a shared classifier to ensure consistent categorical responses across domains. Extensive experiments on challenging benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches.
Domain generalization (DG) attempts to generalize a model trained on single or multiple source domains to an unseen target domain. Motivated by the transferability of vision-language pretrained models, we argue that text can provide complementary semantic cues for domain generalization. In this paper, we develop a Text...
Si-Lei Shen, Jing-Yi Zhang, Yuxi Wang et al.· Electronics· 0 citations
: Lightweight semantic segmentation remains challenging because compact backbones often weaken feature discriminability and lose fine-grained boundary details. In DeepLabV3 + -style encoder-decoder architectures, the direct fusion of high-level semantic features and low-level spatial features may introduce semantic-spa...
Wang Zhang, Lanlan Li, Jiayi Xing et al.· Computers, Materials & C...· 0 citations
Considering the difficulty of learning spatially and semantically aware prompt injection, the Hierarchical Prompt Injector is proposed, which enables spatially adaptive prompt injection in foundation models and auxiliary supervision to align hierarchical prompts with their corresponding object regions is introduced.
Xin Lin, Ruo-Yu Guo, Jia-Qi Guo et al.· 0 citations
Federated parameter-efficient fine-tuning enables distributed clients to adapt pretrained vision-language models without sharing raw data or updating the full backbone. Its effectiveness, however, is limited by domain heterogeneity across clients. Existing personalized methods separate globally shared knowledge from cl...
Wen-Tao Yue, Qing-Yu Mao, Tian-You Lai et al.· 0 citations
Vision-language models show promise in zero-shot semantic segmentation, but a key challenge is the disconnect between text and visual features. While text embeddings can roughly localize unseen objects, they often lack the fine-grained detail necessary for accurate segmentation, leading to oversegmentation or undersegm...
Jia-Xiang Fang, Shi-Qiang Ma, Jing Wang et al.· Neural Networks· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.