Skip to content
Preprint

CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?

Aug 2026 · 0 citations · 31 references
Computer Science

TL;DR

Inspired by human visual change perception, CogVis is proposed, a cognitive memory-guided framework that reformulates OVCD as a perception-memory-verification paradigm that achieves state-of-the-art performance across all evaluated datasets.

Abstract

Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories. Open-Vocabulary Change Detection (OVCD) addresses this need. However, existing methods often entangle temporal perception, semantic discrimination, and region verification, causing unstable results and redundant computation. Inspired by human visual change perception, we propose CogVis, a cognitive memory-guided framework that reformulates OVCD as a perception-memory-verification paradigm. CogVis first employs a Scene Change Perceptron (SCP) to extract a reusable, category-agnostic change prior from frozen bi-temporal features, thereby decoupling temporal evidence from semantic category decisions. A Semantic Memory Calibrator (SMC) then compensates for category-dependent score shifts by dynamically estimating an image-query-specific decision threshold. Finally, an Adaptive Region Filter (ARF) filters connected candidates using learned semantic, temporal, and structural reliability. Experiments on seven benchmarks spanning semantic change detection, binary change localization, and building-damage assessment show that CogVis achieves state-of-the-art performance across all evaluated datasets. By sharing scene-level change perception, CogVis further avoids repeating category-agnostic temporal perception across queries and improves inference throughput by 28.50%.

View source

Similar papers

2026

From Recognition to Reasoning: Edge-Frequency Chain-of-Thought for Remote Sensing Object Detection

The main bottleneck of current object detectors does not lie in weak feature representation. It lies in the lack of structured reasoning. Most detectors follow a recognition-style pipeline. They map features directly to detection results. This design ignores structural relations inside visual information. As a result,...

Xun Li, Yuzhen Zhao, Baoxi Yuan et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Open-World Semantic Segmentation with Sensitivity Modeling

This work addresses open-world semantic segmentation, the joint task of segmenting known classes while detecting and grouping novel or anomalous content without additional supervision, by extending a dual-decoder baseline with a third, complementary decoder within a unified encoder-decoder design.

Anastasios Romanos Varvarigos, Nikos Giakoumoglou, Tania Stathaki · 0 citations
Preprint Aug 2026

VOLA: Improving Open-World Driving by VLM-Based Semantic Attribute Prediction

Driving in the real world is open-world: a car may encounter a fallen mattress, a deer, or other objects outside its training data. Naming them is not enough. The system must know how to treat each region: can it drive over it, and how severe would a collision be? We therefore shift scene perception from category label...

Yuchen Zhang, Yuan Gao, Sebastian Schmidt et al. · 0 citations
Jul 2026

Multimodal Semantic-Probabilistic Objectness for Open World Object Detection

Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes. PROB improves unknown discovery by modeling class-agnostic probabilistic objectness in the decoder-query space. However, visual objectne...

Wei-Jun Tian, Ruining Liu · 0 citations
#machine learning Preprint Aug 2026

Towards Continual Test-Time Adaptation of Vision-Language Models in Open-Vocabulary Semantic Segmentation

Diversify, Anchor, and Filter (DAF), a stabilization framework that augments entropy-based adaptation with a marginal diversity loss that resists collapse, a cross-modal anchor consistency loss that constrains feature drift relative to a frozen source model, and feature salience filtering that skips low-value backward...

Chandler Timm C. Doloriel, Yunbei Zhang, Sarthak Kumar Maharana et al. · 0 citations
Preprint Aug 2026

Continual Visual Learning under Evolving Semantic Concept Shift

Experiments show that SemReWrite achieves a stronger balance between learning revised semantics and retaining unaffected knowledge than prompt replacement, conventional fine-tuning, parameter-efficient adaptation, and continual-learning strategies.

Ismail Lamaakal, Chaymae Yahyati, Yassine Maleh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.