Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Oct 2026

MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs

Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units,...

Xu-Dong Wang, Hao Wu, Hao-Zhe Hu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Supervising Sound Localization by In-the-wild Egomotion

We present a method for learning binaural sound localization using egomotion as a supervisory signal. Over the course of a video, the cameras direction to a sound source will change as the camera moves. We train an audio model to predict sound directions that are consistent with visual estimates of camera motion, which...

Anna Min, Ziyang Chen, Hang Zhao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

PickMoment: Continuous-Time Single-Image-to-Video via Learning Deblurring and Blur-to-Video

Motion blur arises from the temporal integration of a continuous sharp signal over a finite exposure window, yet existing learning-based methods sidestep this physical model and predict only the sharp signal itself: most single-image deblurring methods recover a single frame at the exposure center, while blur-to-video...

Junseong Shin, Hyeonsu Jo, Daehyun Kim et al. · 0 citations
#artificial intelligence Preprint Oct 2026

When the Judge Acts: Auditing VLM-Guided Image Selection on Culturally Situated Prompts

Vision-language models (VLMs) increasingly act as judges that pick the best of several generated images, so their choices decide what users see. Such judges are usually validated by score agreement with human ratings, not by the images they return. We audit VLM judges as decision-makers: on 300 culturally situated prom...

Huichan Seo · 0 citations
#artificial intelligence Preprint Open access Oct 2026

CineMR: Tool-Integrated Vision-Language Reasoning for Quantitative Cardiac MRI Assessment

Cardiovascular magnetic resonance (CMR), including cine imaging, is a reference standard for the noninvasive assessment of cardiac morphology and ventricular function. Cine CMR interpretation integrates qualitative visual assessment with quantitative measurements of ventricular volumes, ejection fraction, myocardial ma...

Kunyang Li, Hai Nguyen, Joshua Lowe et al. · 0 citations
#artificial intelligence Preprint Oct 2026

RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation

Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial relation segmentation...

Minsu Kim, Jaesung Choe, J. Lee et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Two Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold

An answer candidate in a masked diffusion MLLM can stabilize while its rationale is still unfolding. We distinguish retrospective stabilization of the logged candidate from token commitment, and examine these two clocks relative to rationale generation. Analyzing our results across three visual question-answering bench...

Keuntae Kim, Yong Suk Choi · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions

Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with vision--language model (VLM) captioners replacing sparse alt-text by dense descriptions. A recaptioned corpus is a supervision distribution induced by a documented captioning policy ($\pi$), captioner ($V_c$), and source corpus ($C$)...

Giyeong Oh, Junghun Park, Yuhan Bae et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Kinematic MeanFlow: One-Step Action Generation Policy for Robotic Foundation Models

In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a promising framework for this goal, yet its direct application leads to performance collapse. We discover that this stems f...

Jiawei Fan, Sifeng Wang, Yuqing Hou et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Don't Waste the Noise: Importance-Guided Perturbation Allocation under Joint Global and Local Constraints

Adversarial optimization under a shared $\ell_1$ budget requires deciding not only how much perturbation to use, but also where that limited budget should be spent. This allocation problem becomes particularly important when individual input coordinates are subject to local magnitude constraints, which restrict the ext...

M. Shirian, Kianoosh Vadaei · 0 citations
#artificial intelligence Preprint Oct 2026

Geometric Similarity in VLM Low-Level Vision Representations

Vision-language models (VLMs) have emerged as powerful candidates for universal vision backbones, with representative architectures including autoregressive (AR) models and diffusion transformers (DiTs). Yet, adapting them efficiently for all-in-one low-level image restoration remains a challenge. Crucially, the field...

Shao-Jun Xia, Hui-Xin Zhang, Zhen Lei et al. · 0 citations
#artificial intelligence Review Sep 2026

Video Generation Models: A Survey of Post-Training and Alignment

Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maintain temporal coherenc...

Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni et al. · 3 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.