Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to local details. However, most Multimodal Large Language Models (MLLMs) rely on only the final or penultimate vision encoder representations or fixed aggregation rules, making...
Jeonghwan Kim, S. Stoica, Ji-Wan Chung et al.· 0 citations
Despite rapid progress in automating scientific research, generating promising and well grounded research solutions remains a central challenge. We isolate research ideation as a standalone task and build our solution on the intuition that a challenge in one field can often be addressed by a mechanism that solved an an...
Jia-Rui Liu, Ren-Jie Tao, Yi-Wei Liao et al.· 0 citations
Long-term egocentric video enables personalized AI assistants to reason about daily life. However, as video histories grow to hundreds of hours spanning months or years, reprocessing raw clips for every query becomes computationally prohibitive. Memory systems offer a scalable alternative by compacting videos into text...
Guang-Zhi Xiong, Xin-Yuan Zhang, Xiao Yang et al.· 0 citations
This work introduces S-EMBER (Streaming Egocentric Memory Benchmark for Episodic Retrieval), a large-scale benchmark comprising 3,141 videos totaling 388 hours of organic activity captured via Ray-Ban Meta smart glasses that establishes a hardware-authentic foundation for developing grounded, reliable episodic memory i...
Xiaodong Wang, Xuanyi Zhao, Pedro Rodriguez et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.