Energy-Shaped Visual Prompting (ES-VP), a novel approach that generates image-specific prompts using low-rank initialization and energy-guided dynamic adaptation, achieving superior performance with fewer parameters compared to single-prompt methods, thereby establishing a new benchmark for efficient and generalizable model adaptation.
Abstract
Visual prompting (VP) has emerged as a parameter-efficient method for adapting pre-trained models to downstream tasks. However, existing approaches encounter a trade-off between flexibility and efficiency. Some methods apply a fixed prompt to all images, ignoring individual image characteristics, while others introduce auxiliary networks to generate diverse prompts. Although the latter can improve performance, it also significantly increases parameter usage and the potential for overfitting to specific datasets. Furthermore, the auxiliary networks, combined with inherent biases in pre-trained models, limit scalability and generalization. In this paper, we propose Energy-Shaped Visual Prompting (ES-VP), a novel approach that generates image-specific prompts using low-rank initialization and energy-guided dynamic adaptation, achieving superior performance with fewer parameters compared to single-prompt methods. ES-VP directly utilizes the pre-trained model for adaptive prompt generation, ensuring both parameter efficiency and improved generalization. Extensive experiments conducted on five architectures across fifteen datasets demonstrate that ES-VP consistently outperforms current state-of-the-art (SOTA) single and diverse VP methods. For instance, using the CLIP architecture across four datasets, ES-VP outperforms the SOTA method DAM-VP by an average of 2.6\% in accuracy while utilizing 590$\times$ fewer VP parameters, thereby establishing a new benchmark for efficient and generalizable model adaptation.
This work injects two complementary semantic priors into Visual prompt tuning, a cascaded scheme that integrates both priors throughout ViT adaptation, and proposes a cascaded scheme that integrates both priors throughout ViT adaptation.
Xianye Xiao, Xingjian Li, Cheng Han et al.· Trans. Mach. Learn. Res.· 0 citations
Vision-language models (VLMs), such as contrastive language-image pre-training (CLIP), exhibit powerful zero-shot generalization capabilities. Parameter-efficient fine-tuning (PEFT) techniques, notably prompt learning, have been extensively explored to adapt these models to downstream tasks. However, their efficacy remains constrained when transferred to specialized domains like remote sensing. We argue that the bottleneck stems not merely from the limited parameters of prompts, but essentially from the disruption of the input’s original image–text features and the lack of deep cross-modal alignment. In particular, existing methods typically rely on global attention or coarse-grained feature mapping. This inadvertently corrupts the original input representations, thereby impairing the model’s inherent generalization. Furthermore, their isolated unimodal gradient updates fail to bridge the semantic gap inherent in complex remote sensing scenes. To address these challenges, we propose tokenwise prompt-free learning (Tiper), shifting the optimization paradigm from introducing external prompts to precisely recalibrating the critical tokens that govern classification outputs. In particular, Tiper employs a hierarchical learner to supersede global prompts. Crucially, this learner intervenes exclusively on the specific core tokens (i.e., the CLS token in the visual branch and the EOT token in the textual branch), leaving other original input representations unperturbed. This fine-grained strategy effectively balances domain adaptation with the preservation of inherent generalization. Finally, we design the learner as a cross-modal coupled bridge with shared weights, enabling it to synchronously receive gradient feedback from both modalities and fostering profound multimodal collaboration. Extensive experiments validate our method on eight public remote sensing datasets covering diverse scenes and resolutions. In the base-to-new generalization task, Tiper outperforms the strong baseline MaPLe with a significant 3.7% improvement in the harmonic mean (HM). Notably, without relying on any external large-scale domain models, Tiper surpasses the latest domain-specific prompt learning methods (e.g., domain-controlled prompt learning (DCPL), domain prompt learning with quaternion networks (DPLQ)), demonstrating its superior adaptability for remote sensing image scene classification.
Tengfei Gong, Jun-Lin Wu, Yaxioong Chen et al.· IEEE Transactions on Geoscie...· 0 citations
The core of SSVAL is Visual Anchor Prompt Injection (VAPI), which introduces prompts that absorb rich knowledge from external VFMs during training, enabling them to serve as stable visual anchors that mitigate representation deviation during inference.
Qian-Long Yang, Bowen Ye, Xianda Guo et al.· 0 citations
Generalized zero-shot learning (GZSL) addresses the challenging task of recognizing both seen and unseen classes by leveraging shared semantic knowledge. A core challenge in this domain is achieving robust visual-semantic alignment to transfer knowledge from seen classes to novel classes. Current state-of-the-art methods typically fine-tune large-scale visual backbones on scarce training data. However, this approach frequently leads to severe overfitting to seen classes, which significantly degrades performance on novel categories. To mitigate this issue, we propose the Dual Adaptive Visual-Semantic Prompt Collaboration Network (VSPCN+), a novel framework that utilizes prompt-tuning for effective feature adaptation. Our method introduces a dual-prompt mechanism comprising both visual and semantic prompts. The semantic prompts guide the visual encoder to learn visual features that are more semantically consistent with class attributes, while the visual prompts steer the semantic encoder to generate semantic representations that are more visually grounded. This collaborative process enhances the overall visual-semantic consistency. A key innovation of our work is the dynamic generation of instance-adaptive prompts, which contrasts with existing prompt-learning methods that rely on static, global prompts. By tailoring prompts to individual instances, our approach enhances the model’s robustness and generalization capabilities across diverse visual inputs. This collaborative adaptation, guided by our dual-prompt mechanism, allows the visual and semantic encoders to produce consistent representations for effective visual-semantic alignment. Extensive experiments on standard GZSL benchmarks demonstrate that our proposed VSPCN+ performs favorably against several state-of-the-art methods.
Huajie Jiang, Zheng-Xian Li, Yuankai Qi et al.· International Journal of Com...· 0 citations
In CoDA, a new adaptation framework that explicitly disentangles and coordinates cross-modal semantic alignment and intra-modal structural consistency is proposed, and it is shown that CoDA outperforms state-of-the-art parameter-efficient methods, particularly under few-shot learning and distribution-shift scenarios.
Continual learning faces the persistent challenge of catastrophic forgetting, where sequential task updates degrade previously acquired knowledge. While prompt-based methods integrated with pre-trained models offer a compelling solution by freezing the backbone, they often rely on static, task-level prompting strategies that overlook fine-grained intra-task diversity. In this paper, we propose Gated Adaptive Prompting (GAP-Prompt), a novel method that introduces instance-level adaptability to the prompting process. GAP-Prompt consists of three synergistic modules: (1) instance-conditioned gating, which dynamically determines optimal prompt injection layers for each individual image; (2) dynamic knowledge fusion, which performs instance-aware aggregation of current and historical prompts, enabling knowledge integration across tasks; and (3) shared prompt distillation, which anchors foundational knowledge in early shared layers to mitigate forgetting. Extensive evaluations on CIFAR-100, ImageNet-R, and CUB-200 benchmarks demonstrate that GAP-Prompt consistently achieves state-of-the-art performance. Notably, on the fine-grained CUB-200 dataset, GAP-Prompt reaches 87.29% accuracy, approaching the joint training upper bound (88.00%) and outperforming existing methods by a significant margin.
Trung-Anh Dang, D. Bùi, Ngoc-Son Vu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.