Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

Weave Mamba Fusion: Global Cross-Scale Interaction for Lightweight Face Detection

Feature pyramid methods, from FPN to BiFPN, have achieved strong performance in face detection by fusing multi-scale features. However, detecting faces under unconstrained conditions, such as small scale, occlusion, and extreme pose, remains difficult, as it requires global cross-scale dependencies that local fusion ca...

Dohun Kim, Jinmyung Jung · 0 citations
#artificial intelligence Preprint Oct 2026

Imagine to Act: High-Fidelity Data Synthesis via Image Editing World Model for Scalable GUI Agent Training

Graphical User Interface (GUI) agents have emerged as a promising paradigm for automating complex digital workflows across diverse applications. However, training highly capable and generalizable agents fundamentally relies on massive, high-fidelity visual-action trajectories, which are notoriously difficult to acquire...

Yong-Xin Ning, Run-Liang Niu, Qianli Xing et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Level-of-Token Diffusion

Image and video diffusion models allocate equal computation to every region, even when the intended scene calls for varying levels of detail. The spatial distribution of detail can often be anticipated before generation, indicating where computation can be reduced. We introduce Level-of-Token (LoT) Diffusion, a framewo...

Kiyohiro Nakayama, Brian Chao, Jan Ackermann et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models

Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-size states with substantially lower memory cost. However, we...

Zhuokun Chen, Feng Chen, Xi Lin et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Rotated, but How Far? Diagnosing and Improving Object-Rotation Reasoning in VLMs

Vision-language models (VLMs) can detect that an object has rotated across views, but cannot reliably tell by how much. We introduce OR-Bench, a fine-grained benchmark for object-rotation reasoning with eight tasks covering rotation detection, rotation magnitude estimation, and multi-view rotation reasoning. Across 12...

Zhaochen Wang, Yujun Cai, Huangbo Zou et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text...

Team Kandinsky, Julia Agafonova, Bulat Akhmatov et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Human-Like Attention? A Psychophysical Comparison of Visual Search in Humans and MLLMs

Visual search is a fundamental cognitive ability. This study investigates whether Multimodal Large Language Models (MLLMs) exhibit human-like difficulty signatures in visual search tasks. We compared search performance of humans (n = 1,250) and MLLMs using identical 2D and 3D stimuli across different set sizes. Both gr...

Renchi Zhang, Joost C. F. de Winter, Dimitra Dodou et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Strong Baseline for Evaluating Vision Encoders in Multimodal Large Language Models

Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain the dominant choice in practice. In this work, we revisit c...

Yilin Yang, Jun-Tao Tang, Kengyi Wang et al. · 0 citations
#artificial intelligence Preprint Oct 2026

IRSTD-Agent: Agentic Infrared Small Target Detection via Zoom-Guided Interaction Learning

Infrared small-target detection plays an important role in maritime monitoring and aerial surveillance. Although multimodal large language models (MLLMs) offer promising capabilities for visual understanding, existing MLLM-based approaches struggle to precisely localize infrared small targets. In this paper, we propose...

Jia-Wen Xi, Yu Zhang, Tian-Yi Zhao et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Revisiting Ground-Truth Synthesis from High-Speed Video: Exact Validity Conditions and an Audited Consumer Capture Corpus

Motion-deblurring datasets are commonly synthesised by averaging $N$ consecutive high-frame-rate frames and labelling the result with the central frame. We show that this label is unbiased for every capture timing only when the window is odd and the frames' sample durations are equal. An even window shifts every label...

Abdullah Al Shafi, Sumaiya Rahim Suma · 0 citations
#artificial intelligence Preprint Oct 2026

MGPO: Manifold-Guided Diffusion Alignment for Task-Aware Dataset Distillation

Diffusion-based dataset distillation (DD) suffers from a fundamental objective mismatch: likelihood-driven diffusion models prioritize density approximation over the discriminative decision boundaries required for downstream tasks. Beyond semantic mismatch, relying solely on density also leads to geometric coverage los...

Yun-Yi Chen, Chenru Wang, Xin-Yi Ye et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

How Does Geometry Enter Generated Motion?

Under a fixed physical law, the visible geometry of a scene determines how motion must change. We ask how video generators realize this relationship. We fix the law and the initial state and change only the geometry drawn in the first frame, within matched families of tracks and deflectors, and compare each generated t...

Weihan Li, Junhao Wu, Yuhan Song et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.