A dynamic schema-guided world model optimized for visual dynamic prediction and simulation, DynaVieW, achieves an in-depth understanding of visual dynamics by learning interleaved state-transition sequences, where states cover broad visual scenes from video keyframes and transitions capture comprehensive dynamic constituents within a hierarchical schema.
VideoTreeSearch (VTS) is proposed, a framework that casts grounded LVQA as iterative self-correcting search over an adaptive temporal tree, and trains an agent to navigate the tree through four discrete operations: zoom_in, zoom_out, shift, and answer.
Ce Zhang, Ziyang Wang, Yu-Lu Pan et al.· arXiv.org· 0 citations
Compared to prior CLIP-enhancement methods, MLLMCLIP achieves state-of-the-art compositional accuracy while delivering consistent gains on standard zero-shot classification and image-text retrieval, showing that feature-level distillation strengthens both compositional and general vision-language representation capability.
Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.