Skip to content
Open access

MDRAN: a multi-source dual-refining adaptation network for source-free domain adaptation guided by vision-language-model

Aug 2026 · Journal of King Saud University: Computer and Information Sciences · Vol 38 · 0 citations · 49 references

TL;DR

This work proposes a novel criterion, termed Maximum Refinement Decision score (MRD-score), which replaces softmax with median centering to normalize source models’ predictions along both positive and negative axes, thereby harnessing both affirmative and complementary guidance.

Abstract

Source-Free Domain Adaptation (SFDA) adapts a pre-trained source model to an unlabeled target domain without accessing source data, alleviating the need for direct access to source data during adaptation, which reduces the risk of data transmission. Existing methods leverage pre-trained Vision-Language (ViL) models to provide supervisory signals, but their predictions are often noisy, limiting adaptation performance. Moreover, current approaches refine ViL predictions only within individual source domains, overlooking the complementary discriminative evidence available across multiple sources. To better select the most decisive refinement signal from each source model, we propose a novel criterion, termed Maximum Refinement Decision score (MRD-score), which replaces softmax with median centering to normalize source models’ predictions along both positive and negative axes, thereby harnessing both affirmative and complementary guidance. Building upon MRD-score, we present Multi-source Dual-Refining Adaptation Network (MDRAN), which, to the best of our knowledge, is among the first attempts to exploit ViL-based guidance for multi-source-free domain adaptation (MSFDA). MDRAN first aggregates the MRD-score-based predictions from multiple source models to construct Selected Source-Logit Set (SLS), which is then used to customize a collective ViL model via prompt learning. To mitigate noise in the collective ViL models’ outputs, we design a Dual-View Refinement (DVR) module that refines ViL model’s predictions using source model predictions. Extensive experiments on four benchmarks demonstrate that our method achieves state-of-the-art performance on the more challenging Office-Home, VisDA-C, and DomainNet-126, while remaining highly competitive on the comparatively easy Office-31. The code can be found in https://anonymous.4open.science/r/MDRAN-098E/.

Read PDF

Similar papers

Open access Aug 2026

Towards source-free domain adaptation: a model-heterogeneous perspective

A Peer-level Heterogeneous Perception Framework is proposed that departs from such paradigms by enabling balanced collaboration between heterogeneous models by introducing an auxiliary domain that is significantly different from the target domain and employ an auxiliary model with the same architecture as the source mo...

Zhi-Ze Wu, Yu-Tao Fu, Huan-Xin Zou et al. · 0 citations
Open access Sep 2026

A novel source-free domain adaptation with high-confidence sample selection and feature disentanglement for machinery fault diagnosis

This paper proposes a novel SFDA with high-confidence sample selection and feature disentanglement for machinery fault diagnosis, which effectively alleviates the adverse influence of noisy pseudo-labels during the stage of adaptation.

Yi-Ming Yuan, Kang Wu, Xing-Xing Jiang et al. · 0 citations
Aug 2026

Active Domain Adaptation Under Concept Shift.

This paper proposes ADA-CS, a plug-and-play module compatible with any ADA or ASFDA framework, and introduces a CSS metric to quantify the Concept Shift Severity across domains, revealing that non-negligible concept shift exists in many transfer tasks.

Zi-Kang Zhu, Yiyan Huang, Xing Yan · 1 citation
Preprint Aug 2026

Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models

PuRF is introduced, a novel PuRiFication-driven cache-based method for multi-label test-time adaptation of vision-language models that consistently outperforms state-of-the-art methods on ViT-B/32 across five datasets.

Yiwen Liang, Hui Chen, Yizhe Xiong et al. · 0 citations
Preprint Sep 2026

Test-Time Logit Prompting for Source-Free Missing Modality Adaptation

Vision-language models (VLMs) have achieved remarkable performance by leveraging complementary information from large-scale image-text pairs. However, missing-modality inputs are commonly encountered during real-world deployment, often leading to significant performance degradation. Existing methods primarily enhance m...

Taixi Chen, Nancy L. Guo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.