Long video understanding relies on video memory to overcome the context limits of multimodal large language models. Existing methods follow a build-then-reasoning pipeline: memory is built offline for the entire video, then reasoned over as a static source. In practice a long video is shared by several questions, and t...
Wei Chen, Xuan-Yu Zheng, Yan-Cheng Long et al.· 0 citations
The Global State Model (GSM), a causal encoder--decoder architecture that concentrates the selection and aggregation of long-range information in the encoding stage, is introduced, offering a shared-state architecture for efficient language modeling.
Yun Zheng, Bin Wen, Xiao-Jie Wang et al.· 0 citations
Concentrate and Concentrate (CaC) is a coarse-to-fine anomaly reward model based on Vision-Language Models that first conducts a global temporal scan to anchor anomalous time windows, then performs fine-grained spatial grounding within the localized interval, and finally derives robust judgments via structured spatiote...
Jiyuan Wang, Huan Ouyang, Jiu-Zhou Lin et al.· arXiv.org· 5 citations
Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Research supports multi-step information seeking, the two capabilities are...
Wenqi Liu, Shijie Ma, Yun-Xiao Wang et al.· 0 citations
Transformers lack a native lookup mechanism, requiring repeated dense computation to recognize and reuse local static patterns. Lngram v1 introduces tokenizer-independent conditional memory through discrete latent n-gram addressing, but its memory capacity is coupled with the backbone width, limiting scalability due to...
Yun Zheng, Bin Wen, Xiao-Jie Wang et al.· 0 citations
MPAR-Bench is introduced, a bilingual English-Chinese benchmark that isolates reasoning breadth through multi-point associative reasoning, and case-level analysis shows that extended reasoning can overturn an initially correct hypothesis.
Si'an Xie, Jiaxu Liu, Biao Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.