Skip to content

TextFM: Robust Semi-dense Feature Matching with Language Guidance

· 2 citations · 48 references

TL;DR

TextFM is presented, the first language-guided feature matching framework that incorporates domain-invariant semantic information from vision-language models (VLMs) to establish a strong foundation for robust and generalizable feature matching under real-world constraints.

View source

Similar papers

Open access Aug 2026

RoFLIP: Robust and Fine-Grained Alignment for Vision-Language Compositional Reasoning

The Robust and Fine-grained training framework for CLIP-based vision-language models (RoFLIP) is proposed, enhancing both the robustness and granularity of vision-language alignment and underscore RoFLIP’s compositional reasoning and generalization abilities.

Yiwei Sun, Chuanbin Liu, Shancheng Fang et al. · 0 citations
Jul 2026

MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning

This work introduces \method, a Multi-scale Adaptive Vision Encoder, a Multi-scale Adaptive Vision Encoder that uses position-dependent gates to fuse shallow, intermediate, and deep features from a vision Transformer, preserving global semantics while enhancing edges, text, and local structure.

Sha Lei · 0 citations
Jul 2026

TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions

Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions. CLIP provides a strong foundation for these tasks by learning a shared image-text embedding space from large-scale contrastive pre-training. However, its image-level objective aligns text with a CLS-derived global representation, leaving local vision-language correspondence only indirectly constrained. Existing methods either introduce additional supervision, external models, or task-specific adaptation, while training-free approaches mainly recover dense responses from existing patch features without examining where local semantics become most accessible within CLIP. We introduce TraceCLIP, a training-free framework that recovers latent patch-level semantic evidence by isolating the patch-specific terms written into the CLS attention output. TraceCLIP further converts contribution-derived semantic responses into a semantic-geodesic topology gate that calibrates final-layer patch affinity for dense feature reconstruction. Diagnostic experiments show that these contribution features exhibit strong local semantic discrimination and text-conditioned spatial alignment. On eight zero-shot semantic segmentation benchmarks, TraceCLIP achieves gains of 1.3 to 4.5 points in average mIoU over the strongest prior training-free methods across both backbones and background settings, without additional training, external vision foundation models, or region-level supervision. More broadly, these findings suggest that spatially localized semantics may remain accessible within the internal construction of globally aligned representations.

Xinran Liu, Shouqian Shi, Yutong Chen et al. · 0 citations
Preprint Aug 2026

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Qwen-3D is introduced, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes and incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene representation.

Lucy Lin, Ayush Jain, Yifan Liu et al. · 0 citations
Open access Aug 2026

Dynamic Uncertainty Learning with Noisy Correspondence for Text-based Person Retrieval

Text-Based Person Retrieval (TBPR) aims to locate a person in an image database based on a natural language description. While effective in theory, TBPR faces substantial challenges in real-world scenarios due to noisy correspondences—misaligned or weakly related image-text pairs—that significantly degrade retrieval performance. Existing methods often overemphasize hard negative mining, which inadvertently magnifies the impact of such noise. To address this issue, we propose Dynamic Uncertainty with Noisy Correspondences (DUNC), a novel framework that incorporates two key components: (1) Cross-modal Evidential Learning (CEL), which models bidirectional alignment uncertainty using a Dirichlet distribution to capture the confidence in image-text similarity, and (2) Dynamic Robust Loss (DRL), which adaptively selects and aggregates hard negative samples to reduce the influence of noisy instances and improve model robustness. Unlike conventional global-alignment approaches, DUNC exploits fine-grained local correspondences to enhance semantic alignment between modalities. By integrating uncertainty-aware modeling and adaptive contrastive supervision, our method is capable of effectively disentangling noisy from reliable training pairs. Extensive experiments conducted on three benchmark datasets—CUHK-PEDES, ICFG-PEDES, and RSTPReid—demonstrate that DUNC consistently achieves state-of-the-art performance and exhibits strong robustness across a wide range of noise conditions. Code is publicly available at https://github.com/ASL-forever/DUNC.

Zequn Xie, Chuxin Wang, Sihang Cai et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.