ViewMind3D is presented, a fully training-free and modular framework for 3D spatial reasoning over multi-view observations of a scene without requiring complete 3D reconstruction, which improves performance on spatially grounded question types, such as ``What''questions in SQA3D, while maintaining strong overall accura...
Ping-Kun Chiang, Kun-Ru Wu, Po-han Li et al.· arXiv.org· 0 citations
This work proposes LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation, and introduces Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences.
Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.