Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Review Oct 2026

VideoResearchAgent: Grounded Task Synthesis and Sim-to-Real RL for Open-Web Video Research

Existing deep research agents are designed primarily for text- and image-based web sources, while video reasoning systems typically assume that relevant videos are provided in advance. We study open-web video research, where an agent must autonomously discover relevant videos, navigate their temporal content, and groun...

Yu-Hang Zhou, Fei-Yu Li, Yu-Xin Wu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

MedImageOSWorld: Benchmarking GUI Agents for Medical Image Consoles

Graphical consoles offer a practical interface for medical acquisition assistance, allowing agents to work through the controls and visual feedback used by human operators. Reliable assistance requires linking on-screen anatomy to acquisition decisions that determine what image evidence becomes available next. We intro...

Ziyang Long, Xinqi Li, Lujing Xing et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness

Missing modality remains a longstanding challenge in multimodal learning. Existing methods typically address this issue through modality recovery or adaptive strategies. However, they overlook models'internal cross-modal dependencies formed during multimodal training, which later impair robustness. We systematically ch...

Song-Yuan Sui, Zhen Tan, Mohan Zhang et al. · 0 citations
#computer vision Preprint Open access Oct 2026

A Benchmark for Spatially Grounded Gesture Generation

Communication in shared space interweaves verbal and non-verbal signals, and pointing gestures anchor language to the environment: "put the cup on that one" is uninterpretable without the gesture that fixes the referent. Yet no common framework exists for evaluating whether generated gestures indicate their intended re...

Anna Deichler, Rishabh Dabral, Fethiye Irmak Dogan et al. · 0 citations
#computer vision Preprint Open access Oct 2026

Last But Not Least: Boundary Attention CalibratiON for Multimodal KV Cache Compression

Multimodal Large Language Models (MLLMs) achieve strong vision-language reasoning but incur large KV caches and high decoding latency with long visual contexts. Existing compression methods rely on observation window attention for stable token importance estimation, yet this aggregation can dilute sparse critical evide...

Tianhao Chen, Yuheng Wu, Kelu Yao et al. · 0 citations
#computer vision Preprint Open access Oct 2026

WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models

Contrastive vision-language models have achieved remarkable progress through large-scale pretraining. Recent work has shown that removing English-only caption filters and pretraining on global data is effective for improving multicultural performance. We study whether such global pretraining is sufficient for culture-s...

Issa Sugiura, Shuhei Kurita, Yusuke Oda et al. · 0 citations
#computer vision Preprint Open access Oct 2026

The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?

Cognitive science research treats visual perception, the ability to understand and make sense of a visual input, as one of the early developmental signs of intelligence. Its TVPS-4 framework categorizes and tests human perception into seven skills such as visual discrimination, and form constancy. Do Multimodal Large L...

Samrajnee Ghosh, Ashish Goswami, Naman Agarwal et al. · 0 citations
#computer vision Preprint Open access Oct 2026

World Embedding Benchmark

Physical fidelity has received increasing attention in world models and video generation, yet how video representations encode physical information remains less understood. We introduce the World Embedding Benchmark, comprising 8,000 controlled simulation cases from 80 families spanning fluid mechanics, solid mechanics...

Yiqi Liu, Ruifeng Yuan, Yang Wang et al. · 0 citations
#computer vision Preprint Oct 2026

ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation

Simultaneous Sign Language Translation (SLT) is critical for real-time communication, yet existing methods remain largely confined to sentence-level, offline settings that assume pre-segmented inputs. These assumptions hinder deployment in realistic scenarios involving continuous, unsegmented video streams. We present...

Si-Han Ren, Gao-Zheng Li, Yuan-Shang Quan et al. · 0 citations
#computer vision Preprint Open access Oct 2026

Recursive Self-Improvement in Unified Multimodal Models

Unified multimodal models (UMMs) understand and generate both text and images, which lets a model produce its own training data. Existing self-improvement in UMMs keeps supervision on the visual side, where image understanding judges image generation. We propose recursive cross-capability self-improvement (RSI), a trai...

Huijuan Wang, Chufan Shi, Cheng Yang et al. · 0 citations
#machine learning Preprint Open access Oct 2026

GB-LSR: Local Spectral Decoding with a Learned Global Bandwidth for Arbitrary-Scale Super-Resolution

We present GB-LSR (Global-Bandwidth Local Spectral Representation), a fixed-grid local spectral representation for continuous image decoding. The image domain is partitioned into non-overlapping square patches. Each patch carries coefficients for a truncated Fourier basis, predicted by a single linear projection from s...

Max Shad, Naeem Khoshnevis · 0 citations
#machine learning Preprint Open access Oct 2026

Low-Frequency Shortcuts in Texture-Driven Visual Learning

Neural networks suffer from shortcut learning, where learned features generalize well to the training set but not to in-distribution (ID) or out-of-distribution (OOD) test sets. Existing studies are all based on a few standard benchmarks, which are shape-driven. Numerous application domains, however, are texture-driven...

Utku \c{S}irin, Cathy Hou, Despina-Ekaterini Argiropoulos et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.