CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
Results support the effectiveness of caption memory for episodic reasoning over long egocentric video in wearable assistants with bounded frame budgets, growing visual-token costs, and long-context retrieval failures.
Dingli Liang, Yi Xie, Yu-Kai Huang et al.
· 0 citations