This work introduces LAGO (LAnguage-Guided adaptive Object-region focus), which reframes localized recognition as language-guided directed region discovery and addresses the circular dependence between recognizing a class and locating its supporting evidence, while preserving complementary local, contextual, and global...
Jun-Yi Hu, Qiji Zhou, Lei Zhang et al.· arXiv.org· 0 citations
V-Co, a systematic study of visual co-denoising in a unified JiT-based framework, outperforms the underlying pixel-space diffusion baseline and strong prior pixel-diffusion methods while using fewer training epochs, offering practical guidance for future representation-aligned generative models.
Han Lin, Xichen Pan, Zun Wang et al.· arXiv.org· 0 citations
This work introduces DEER-3D, an error-driven framework that diagnoses grounding failures and generates targeted counterfactual training supervision via a structured "Decompose, Diagnose, Edit, and Retrain"loop, and underscores the effectiveness of targeted, error-driven scene editing in bridging linguistic reasoning w...
Yue Zhang, Zun Wang, Han Lin et al.· arXiv.org· 0 citations
EgoMemReason is introduced, a comprehensive benchmark for week-long egocentric video understanding through memory-driven reasoning that evaluates three complementary memory types: entity memory, tracking how object states evolve and change across days; event memory, recalling and ordering activities separated by hours...