Results support the effectiveness of caption memory for episodic reasoning over long egocentric video in wearable assistants with bounded frame budgets, growing visual-token costs, and long-context retrieval failures.
Dingli Liang, Yi Xie, Yu-Kai Huang et al.· 0 citations
This work proposes Caption-once, Frames-onDemand (CFD), a budget-aware edge-cloud agentic framework that turns visual access into a first-class, query-conditioned cost, capping per-query frame consumption regardless of video length.
Wei-Tong Cai, Hang Zhang, Yu-Kai Huang et al.· 2 citations
Neural radiance fields (NeRF) and 3D Gaussian Splatting (3DGS) are popular techniques to reconstruct and render photorealistic images. However, the prerequisite of running Structure-from-Motion (SfM) to get camera poses limits their completeness. Although previous methods can reconstruct a few unposed images, they are...
Yu Chen, Rolandos Alexandros Potamias, Evangelos Ververas et al.· Neural Information Processin...· 0 citations
This work repurposes pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task, and inherits naturally structured knowledge and richer priors from the video model, enabling more data efficient and effective learning of...
Haosen Yang, Ji-Fei Song, Zhensong Zhang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.