Open-Vocabulary Audio-Visual Semantic Segmentation (OV-AVSS) aims to perform pixel-level segmentation of sound-emitting objects from an open set of categories. The previous method relies on a class-agnostic foreground definition, which groups semantically diverse objects into a heterogeneous positive set, causing the model to learn unstable sounding patterns and produce unreliable proposals. To address this, we reformulate the objective to be category-specific and propose a novel Acoustically Grounded Cost Learning (AGCL) framework to transform the static, audio-agnostic visual-text priors into dynamic, audio-grounded cost representations. For intra-category soundingness discovery, we devise Audio-Modulated Cost Generation (AMCG) and Audio-Guided Temporal Aggregation (AGTA) modules to enable both frame-level sounding region highlighting and video-level temporal refinement with a low-intrusive audio injection mechanism. For inter-category distractor discrimination, we introduce a Synergistic Distractor Mining (SDM) strategy, which selectively penalizes acoustically and semantically confusing negative categories to learn more discriminative decision boundaries. Extensive experiments on the AVSBench-OV dataset demonstrate that our method significantly outperforms previous state-of-the-art approaches, particularly on unseen categories. Code is available at https://github.com/spyflying/AGCL.
Tianrui Hui, Shaofei Huang, Qi-Song Han et al.· 0 citations
The proliferation of AI-generated content is fundamentally altering the information ecosystems in which retrieval systems operate. Search engines, recommender systems, and retrieval-augmented generation pipelines increasingly function in mixed information environments where synthetic and human-authored content are tightly interwoven, raising system-level challenges for information retrieval. Key issues include limitations in evaluation validity, as traditional metrics designed for human-authored corpora fail to capture the distinctive properties of AI-generated content; shifts in retrieval behavior and ranking dynamics, as systems may inadvertently favor procedurally generated but weakly grounded information; and challenges to user trust, as assumptions about the provenance and reliability of retrieved human content become more difficult to distinguish from generated content. Rather than focusing only on model-centric performance comparisons, this workshop aims to provide a forum to analyze these implications with an emphasis on reflection, evaluation, and human-centered system design, and to foster community-driven discussion that may inform future evaluation efforts, including potential shared tasks or tracks in venues such as TREC, CLEF, FIRE, or NTCIR.
Ping Liu, Zhedong Zheng, Shane Culpepper et al.· Annual International ACM SIG...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.