Extensive experiments show that ECADet consistently outperforms representative class-agnostic and proposal-based detectors on BTCO-Bench, demonstrating the effectiveness of expanded objectness discovery.
Abstract
Object detection is a fundamental task in visual perception, providing structured region representations for recognition, grounding, reasoning, and interaction. However, existing detection paradigms largely inherit a thing-centric notion of objectness, where detectors are mainly trained to localize discrete and countable object instances. Consequently, many semantically meaningful visual elements, such as sky, road, grassland, water, and sports courts, are often absorbed into the background despite their importance for scene understanding and spatial reasoning. In this paper, we formulate Expanded Class-Agnostic Detection (ECAD), a new setting that aims to discover category-agnostic visual candidates beyond conventional thing-centric objects. To support this setting, we construct BTCO-Bench, a Beyond Thing-Centric Objectness benchmark with category-agnostic box annotations covering both real-world and cross-domain scenarios. We further propose ECADet, a lightweight DETR-based detector built upon a frozen DINOv3 encoder, and introduce Geometry-Aware Expert Regression (GAER) and Prototype-Guided Query Modulation (PGQM) to improve localization and objectness estimation for diverse visual elements, respectively. Extensive experiments show that ECADet consistently outperforms representative class-agnostic and proposal-based detectors on BTCO-Bench, demonstrating the effectiveness of expanded objectness discovery. Code and benchmark will be released.
Object detectors are typically trained under closed-set supervision, where unlabeled regions are implicitly treated as background. Under incomplete annotations, this assumption introduces objectness bias: visually valid but unlabeled objects are used as negatives, tying objectness to the annotated taxonomy rather than...
Dania Batool, Liliana Lo Presti, M. La Cascia et al.· 0 citations
Locating a specific object instance in a cluttered scene using a single reference image and a short description, and reporting when that instance is absent, large vision-language models usually address this task. We ask whether the same capability is available far more cheaply, from representations already learned by a...
K. Gupta, Ahmed Rafi Hasan, Md. Mahfuzur Rahman et al.· 0 citations
A five-axis taxonomy (modality, mechanism, prompting, supervision level, and generalization setting) is introduced to audit the literature across application domains, including microscopy, remote sensing, crowd counting, and agriculture, and formalizes prevailing challenges into six structural contradictions.
Joana Konadu Owusu, S. Sheshappanavar· 0 citations
This work proposes FineHOI, a zero-shot HOI framework that explicitly models interactions from dense patch-level features, and introduces an Adaptive Part-Level Attention module that decomposes humans and objects into semantically coherent parts via unsupervised clustering, and re-weights them based on their interactio...
Francesco Tonini, Lorenzo Vaquero, Mohammad Mahdi Derakhshani et al.· 0 citations
Class-geometry supervision (CGS) is proposed, a general framework that constrains learned prototype or class-representation spaces to preserve visual or semantic class dissimilarities estimated from training data and suggests that relational class geometry is an effective supervisory signal for building calibrated and...
A. Rao, Zhou Chen, Revanth Reddy Palem et al.· 0 citations
Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene understanding. However, existing COPE methods still require labor-intensive recollection of real-world training data for novel object categories, which limits their scal...
Jian Liu, Wei Sun, Zhen-Qi Dai et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.