The core of D2-ScaleAgent is a Verifier agent-driven dynamic routing loop based on the intrinsic difficulty of the query, centered around a continuously updated evidence bank that serves as the agent's dynamic working memory.
Hao Zhang, Longrong Yang, Lunhao Duan et al.· 0 citations
VideoTreeSearch (VTS) is proposed, a framework that casts grounded LVQA as iterative self-correcting search over an adaptive temporal tree, and trains an agent to navigate the tree through four discrete operations: zoom_in, zoom_out, shift, and answer.
Ce Zhang, Ziyang Wang, Yu-Lu Pan et al.· arXiv.org· 0 citations
ST-Evidence is introduced, the first human-verified benchmark for both discriminative and generative pixel-level grounding, and scalable, automated generation pipelines are developed to create ST-Evidence-Instruct, a 160k-scale dataset bridging high-level reasoning with fine-grained grounding.
Shijie Wang, Honglu Zhou, Ziyang Wang et al.· arXiv.org· 0 citations
EgoMemReason is introduced, a comprehensive benchmark for week-long egocentric video understanding through memory-driven reasoning that evaluates three complementary memory types: entity memory, tracking how object states evolve and change across days; event memory, recalling and ordering activities separated by hours or days; and behavior memory, abstracting recurring patterns from sparse, repeated observations over the whole week period.