Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Sep 2026

Codebook-Guided Cross-Modal Knowledge Distillation for Structurally Heterogeneous Features

Cross-modal knowledge distillation transfers knowledge from a teacher modality to a student modality. Existing feature-level alignment methods typically assume that teacher and student features reside in structurally alignable representation spaces. However, this assumption does not hold when cross-modal features are s...

Dae Ung Jo, Jongin Lim, YoungJoon Yoo et al. · 0 citations
#artificial intelligence Preprint Sep 2026

ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression

ResComEmb is proposed, a trainable framework for effective and efficient universal multi-vector multimodal embedding that produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval.

Zi-Jing Cai, Yu-Zhe Wang, Jing-Xian Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

NowcastDiT: Diffusion Transformers are Effective Precipitation Nowcasters

This work develops NowcastDiT and instantiates this flexibility through two complementary adaptations: a dynamics-aware noise prior for temporally coherent forecasts, and end-to-end reinforcement learning with timestep-aware rewards for meteorological skill.

Hao-Ran Xu, Xing-Zhuo Guo, Yu-Chen Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MotionInsight: Diagnosing Object Motion Deficiencies in Generated Videos

Despite rapid progress in video generation models, they still exhibit obvious motion deficiencies, often manifested as incorrect object motion. However, most existing video quality evaluations focus on aesthetic quality or text-video alignment. To address this gap, we study object-centric motion fidelity assessment, ev...

Jia-Hao Zhan, Yong-Rui Ma, Qunliang Xing et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Embedded Bi-Temporal Building Damage Assessment for On-Board Data Reduction

Rapid assessment of building damage after natural disasters is essential to support emergency response. Earth Observation satellites can acquire relevant imagery shortly after an event, but exploitation is limited by uplink and downlink capacity and by ground-processing latency. We address this with a bi-temporal build...

Thomas Goudemant, Benjamin Francesconi, Marjorie Bellizzi et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom

High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual par...

Xijia Tao, Yihua Teng, Xinyu Fu et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Parameterized Stripe Attention for Efficient Video Generation

Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency...

Xingyu Jia, Baole Ai, Ang Wang et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation

Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-related challenges. Teacher forcing pairs clean history with a noisy target, but supervises predictive contextual representations only indirectly through velocity prediction. Me...

Xingtong Ge, Yutong Wang, Lunjie Zhu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

WeLike2Party! In-Context Motion Transfer for Multi-Human Image Animation

Human image animation aims to transfer motion from a driving video to subjects in a reference image. Despite remarkable progress in video generation, achieving high-fidelity animation of multiple interacting subjects remains a challenge. Many existing approaches rely on explicit motion representations such as 2D skelet...

Sangeyl Lee, Seunghyun Shin, S. Park et al. · 0 citations
#artificial intelligence Preprint Sep 2026

On-Policy Visual Evidence Distillation

Reflection on Visual Evidence (ReVuE), an on-policy distillation method for visual agents that translates trajectory-level evidence diagnosis into targeted token-level supervision, guiding students to improve their visual evidence acquisition and reasoning.

Shaohang Wei, Feifan Song, Guangyue Peng et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Dual-Mode Low-Rank Learner with Bridge-Prototype Ensemble for Vision-Language Class-Incremental Learning

Benefiting from transferable visual-textual alignment, CLIP has been widely adopted for class-incremental learning (CIL). However, existing learners either repeatedly update components shared across tasks, leading to knowledge overwriting, or overly isolate new-task updates, hindering the reuse of CLIP's transferable k...

Chiyuan He, Zihuan Qiu, Fanman Meng et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

NesTok: Nested Self-Aligned 1D Tokenizer for Autoregressive Image Generation

One-dimensional (1D) variable-length visual tokenizers enable adaptive compression by varying the number of tokens, allowing downstream autoregressive (AR) models to flexibly trade off generation quality against computational cost using a single tokenizer. However, existing approaches based on nested dropout often fail...

Jiawei Zhang, Shuhao Liu, Rong Huang et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.