Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

Prompting Image Generators for Training-free Primitive Shape Abstraction

Compact primitive abstractions represent 3D shapes with a few geometric primitives while preserving recognizable components. Learned methods depend on their training classes, and optimization-based methods split shapes geometrically rather than into parts. We instead reuse the visual part knowledge of pretrained models...

Gregor Kobsik, Tim Elsner, Leif Kobbelt · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasets

Computer-use agents turn vision-language model (VLM) predictions into executable GUI clicks, so reliable uncertainty estimates are essential for rejection, calibration, miss-severity ranking, and spatial safety regions. Yet evidence on post-hoc uncertainty quantification (UQ) for these agents is fragmented across isola...

Divake Kumar, Sina Tayebati, Devashri Naik et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory

Vision-Language-Action (VLA) models are increasingly deployed across diverse tasks, yet they remain black boxes whose physical interactions can cause irreversible harm, making generalizable and interpretable failure detection essential. We observe that successful and failed rollouts carry systematically different infor...

Jinghan Yang, Yunchao Zhang, Wang Yuan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ATM: Why Latent World Models Can Fail to Plan

Latent world models can achieve accurate latent prediction yet still differ substantially in downstream planning performance. We argue that a key source of this discrepancy lies in the structure of action-induced latent transitions. We formalize action-identifiability through Bayes inverse risk, characterizing how much...

Jiaheng Chen, Tinghe Zhang, Yucheng Xiao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes

Frontier multimodal large language models (MLLMs) have been reported to achieve over 90\% accuracy on fine-grained perception benchmarks. However, such scores do not necessarily imply faithful use of visual evidence. Prior studies have identified three shortcuts that inflate benchmark performance. First, linguistic pri...

Jingru Chen, Yiming Liu, Mingtao Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Concept Unlearning via Cross-Attention Activation Projection for Diffusion Models

Existing closed-form methods for concept unlearning in text-to-image diffusion models typically derive editing directions from fixed text embeddings, which may not fully capture how concepts are expressed across latent states, timesteps, and layers. To capture this variation, we investigate cross-attention activations...

Saemi Moon, Suhyeon Jun, Seoyeon Lee et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Rethinking Cross-Layer Information Routing in Diffusion Transformers

Diffusion Transformers (DiTs) have become a de facto backbone of modern visual generation, and nearly every major axis of their design -- tokenization, attention, conditioning, objectives, and latent autoencoders -- has been extensively revisited. The residual stream that governs how information accumulates across laye...

Chao Xu, Maohua Li, Qirui Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Latent Video Prediction for World Modeling: An Evaluation Uncovering Intriguing Favorable Evidence

Self-supervised video models are increasingly framed as world models, yet they are still evaluated almost entirely on clean video and reported as a final task score, obscuring how their representations behave under the degraded and ambiguous conditions a deployed world model must handle. We present the first systematic...

Ali J Alrasheed, Aryan Yazdan Parast, Basim Azam et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Image AID via continuous-time reinforcement learning

We study image inpainting with generative diffusion models. Existing methods typically either train dedicated task-specific models, or adapt a pretrained diffusion model separately for each masked image at deployment. We introduce a middle-ground model, termed Amortized Inpainting with Diffusion (AID), which keeps a pr...

Yilie Huang, Xun Yu Zhou · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space

As current Multimodal Large Language Models rapidly saturate canonical visual reasoning benchmarks, a key question emerges: do these strong scores genuinely reflect robust visual understanding? We identify a pervasive vulnerability, the Cartesian Shortcut: models frequently discretize the orthogonal grid-based layouts...

Xia Hu, Zhenrui Yue, Brian Potetz et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models

Multimodal reward models have advanced substantially in text and image domains, yet progress in video understanding reward modeling remains severely limited by the lack of robust evaluation benchmarks and high-quality preference data. To address this, we propose a unified framework spanning benchmark design, data const...

Yuancheng Wei, Linli Yao, Lei Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images

Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition, largely because natural image datasets provide limited supervision for low-level visual skills. Can targeted synthetic supervision address these weaknesses without reference images or manua...

Guanyu Zhou, Yida Yin, Wenhao Chai et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.