Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are s...
Dae Ung Jo, Jongin Lim, YoungJoon Yoo et al.· 0 citations
ResComEmb is proposed, a trainable framework for effective and efficient universal multi-vector multimodal embedding that produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval.
Zi-Jing Cai, Yu-Zhe Wang, Jing-Xian Zhu et al.· 0 citations
This work develops NowcastDiT and instantiates this flexibility through two complementary adaptations: a dynamics-aware noise prior for temporally coherent forecasts, and end-to-end reinforcement learning with timestep-aware rewards for meteorological skill.
Hao-Ran Xu, Xing-Zhuo Guo, Yu-Chen Zhang et al.· 0 citations
Despite rapid progress in video generation models, they still exhibit obvious motion deficiencies, often manifested as incorrect object motion. However, most existing video quality evaluations focus on aesthetic quality or text-video alignment. To address this gap, we study object-centric motion fidelity assessment, ev...
Jia-Hao Zhan, Yong-Rui Ma, Qunliang Xing et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire relevant imagery shortly after an event, but exploitation is limited by uplink and downlink capacity and by ground-processing latency. We address this with a bi-temporal build...
Thomas Goudemant, Benjamin Francesconi, Marjorie Bellizzi et al.· 0 citations
High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual par...
Xijia Tao, Yihua Teng, Xinyu Fu et al.· 0 citations
Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency...
Xingyu Jia, Baole Ai, Ang Wang et al.· 0 citations
Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Me...
Xingtong Ge, Yutong Wang, Lunjie Zhu et al.· 0 citations
Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skelet...
Sangeyl Lee, Seunghyun Shin, S. Park et al.· 0 citations
Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents that translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning.
Shaohang Wei, Feifan Song, Guangyue Peng et al.· 0 citations
Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate new-task updates, hindering the reuse of CLIP's transferable k...
Chiyuan He, Zihuan Qiu, Fanman Meng et al.· 0 citations
One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail...
Jiawei Zhang, Shuhao Liu, Rong Huang et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.
Computer scientist, entrepreneur, and philanthropist will collaborate with the MIT Schwarzman College of Computing to advance AI and scientific discovery.