High-resolution 3D generation increasingly relies on voxel latents and multi-stage pipelines that first predict active structure and then synthesize local geometry. While effective, this design fragments continuous surfaces into many local tokens, inflates generation cost, and often weakens topological consistency for...
Tianjiao Yu, Xin-Zhuo Li, Yi-Fan Shen et al.· 0 citations
GraphVid is introduced, a graph-conditioned image-to-video generation model that enables interactive control through structured interaction graphs that delivers strong controllability and video quality and highlights the potential of structured semantic interfaces as a powerful paradigm for controllable video generatio...
Vedant r Shah, O. Susladkar, Tushar Prakash et al.· arXiv.org· 0 citations
Video editing spans diverse editing paradigms, yet achieving high-quality instruction-guided and subject-guided editing within a single unified framework remains challenging. We introduce EditVid, a training-free framework combining sparse causal memory for local coherence, correspondence-based post-attention token inj...
A. Juvekar, O. Susladkar, K. A. Nguyen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.