Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Rece...

Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc et al. · 0 citations
#artificial intelligence Preprint Oct 2026

GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning

Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly s...

Ya-Kun Zhu, Yi Bin, Yu-Juan Ding et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Task-Adaptive Grounded 3D-Programmers Using 2D VLMs

Recent vision-language models (VLMs) exhibit remarkable generalization and reasoning abilities, yet 3D understanding in these models is limited by data scale, training diversity, and reasoning capacity. Instead of naively extending these models into 3D, we take a different approach: we enable powerful 2D VLMs to operat...

Arman Raayatsanati, Sombit Dey, Anna-Maria Halacheva et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Exploring Weaknesses of Generative Image Watermarks against Latent Frequency Masking

Invisible watermarking has become a central tool for tracing AI-generated images, but its robustness against adaptive removal attacks remains an open security question. We introduce Latent Frequency Masking, an attack that erases watermark evidence by replacing selected Fourier coefficients in the latent representation...

Kirill Aistov, Khaled Abud, I. Serzhenko et al. · 0 citations
#artificial intelligence Preprint Oct 2026

MoLE: Mixture of Latent Experts for Complementary Visual Reasoning

Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value project...

Yingcheng Liu, Tian-Yi Jiang, Yu-Juan Ding et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Unsupervised Domain Adaptation for Enhanced Radiometer Image Precipitation Estimation using Conditional Flow Matching

Deep generative networks have recently achieved unprecedented performance in precise image and video editing using sophisticated textual prompts. However, the effectiveness of such models heavily depends on access to very large supervised and annotated image datasets, which can be very difficult to obtain. This is part...

Victor Enescu, Assaad Zeghina, Matthieu Meignin et al. · 0 citations
#artificial intelligence Preprint Oct 2026

VETO: Video Efficient Token Optimization for Vision Language Models

Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We...

Gueter Josmy Faure, Hao-Ping Wang, Min-Hung Chen et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Cog-VADU: A Training-Free Cognitive Reasoning Framework for Video Anomaly Detection and Understanding

Video Anomaly Detection (VAD) aims to temporally localize abnormal events in videos. Most existing approaches rely on dataset-specific training and curated annotations, limiting generalization in open-set scenarios. Recent zero-shot methods based on Large Vision- Language Models (LVLMs) alleviate this dependency but of...

Mohd Ubaid Wani, Sara Atito, Josef Kittler et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Architectural Sampling: Test-Time Scaling via Computational Diversity in Frozen Vision-Language Models

Test-time scaling often seeks better answers by sampling multiple responses from a frozen model, yet conventional temperature sampling generates every candidate along the same fixed computation path. We introduce architectural sampling, a training-free method that generates candidates through distinct forward computati...

Akshit Singh, Shyam Marjit, Wei Lin et al. · 1 citation
#artificial intelligence Preprint Oct 2026

Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference

Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution...

Xinye Zhao, Yun-Kai Dang, Yun-Chen Wu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Hob-VL: A Benchmark for Visually Grounded Boolean Reasoning

Reliable visual reasoning requires composing multiple visual observations and returning consistent answers to logically equivalent questions. We introduce Hob-VL, a benchmark for visually grounded Boolean reasoning. Hob-VL comprises two tasks: (1) evaluating whether a Boolean rule holds in an image, and (2) identifying...

Yuzhou Wang, Emile Anand, Ijay Narang · 0 citations
#artificial intelligence Preprint Oct 2026

Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs

Video Large Language Models (VideoLLMs) receive frames in sequential order and interpret how visual content evolves along the temporal axis, yet temporal reasoning remains a persistent weakness across architectures. Reversing the frame order of a video, a transformation that should invert temporal answers, often leaves...

Youngwoo Shin, Yusung Ro, Minseo Kim et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.