The Dual-Objective Triggers (DOT) is presented, the first transferable attack on OV-VIS that simultaneously exploits the vision–language coupling and temporal coherence, and Phase-Guided Ad-versarial Training is introduced, which injects perturbations primarily in the phase spectrum while blending amplitudes with clean references.
While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs remain vulnerable to adversarial attacks, posing significant security risks. Existing defense methods predominantly target single-task scenarios (e.g., zero-shot classification) and consequently lack generalizability across various multimodal tasks. To address this limitation, we propose a dual adversarial fine-tuning framework that jointly optimizes visual and semantic supervision signals from two modalities, enhancing model robustness while generalizing across multiple downstream tasks. The proposed framework comprises two core components, i.e., $\textbf{Visual}$ supervision branch and $\textbf{Semantic}$ supervision branch. The former branch leverages features from clean images, extracted via a frozen original vision encoder, to guide adversarial robustness while the latter incorporates caption-image alignment as a contextual signal to preserve semantic coherence under attack. Moreover, our method achieves cross-task robustness by simply replacing the CLIP vision encoder in the original model, with no need of separate task-specific retraining or architecture modifications.Extensive experiments demonstrate that our approach outperforms the state-of-the-art method in adversarial robustness evaluation across zero-shot classification, image captioning, and visual question answering (VQA) tasks.
Sibo Wang, Jie Zhang, Shiguang Shan et al.· arXiv.org· 0 citations
UniVVT is presented, a unified end-to-end framework that reframes VVT as semantically conditioned video generation, eliminating mask, pose, and warping modules at inference and validating implicit semantic guidance as a simple and effective alternative to fragile geometric preprocessing for end-to-end virtual try-on.
Yushe Cao, Shikun Feng, Fei Shen et al.· 0 citations
ST-Evidence is introduced, the first human-verified benchmark for both discriminative and generative pixel-level grounding, and scalable, automated generation pipelines are developed to create ST-Evidence-Instruct, a 160k-scale dataset bridging high-level reasoning with fine-grained grounding.
Shijie Wang, Honglu Zhou, Ziyang Wang et al.· arXiv.org· 0 citations
This work proposes RITA, a Robust test-tIme prompt-TAdaptation framework that shifts from sample-level estimates to distribution-level alignment, and employs optimal transport to align the distribution of augmented visual features with textual prototypes, mitigating adversarial outliers and rectifying cross-modal semantic misalignment.
Xingyu Zhu, Huanshen Wu, Shuo Wang et al.· 1 citation
While generative models have become a standard approach for addressing the semantic-to-visual gap in Generalized Zero-Shot Learning (GZSL), existing architectures often struggle with two persistent limitations: cross-modal interference during condition fusion and severe overfitting to the visual distributions of seen classes. To address these bottlenecks, this paper introduces SemanticFlowNet, a framework based on Decoupled Semantic Flow Matching. Specifically, we propose a Decoupled Multi-modal Conditioning mechanism that relies on channel-wise concatenation of temporal encodings, semantic attributes, and visual contexts, which preserves the orthogonal subspaces of each modality and reduces interference. Additionally, we integrate a Dropout-enhanced Adaptive Layer Normalization (AdaLN) module to perturb the rigid memorization of seen classes, utilizing stochastic dropout within the state evolution to simulate the distributional variance of unseen domains. Finally, a Time-Aware Dynamic Reconstruction Penalty is introduced to enforce progressively stricter semantic alignment as the generative ordinary differential equation (ODE) trajectory converges to the target manifold. Evaluations on the CUB, SUN, and AWA2 benchmarks demonstrate the effectiveness of the proposed framework. Notably, SemanticFlowNet achieves a harmonic mean of 77.90% on the CUB dataset in the single-seed full-model setting, providing a competitive baseline for generative GZSL applications.
Chuyang Song, Mingyi Song, Yang Liu et al.· Electronics· 0 citations
This paper identifies three previously overlooked issues caused by inappropriate cross-modal interactions and excessive operations in the Simple Vision-Language Attack (SimVLA) pipeline, and proposes the SimVLA, which observably improves transferability and efficiency.
Yuchen Ren, Zhengyu Zhao, Chenhao Lin et al.· IEEE Transactions on Informa...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.