Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Sep 2026

Investigating Single-Block Recurrence in Vision Transformers for Image Recognition

Vision Transformers (ViTs) implement depth by stacking independently parameterized blocks, but it remains unclear how much of this parameterization is necessary and how much can be replaced by recurrent reuse. We study this question with bViT, a single-block recurrent ViT that repeatedly applies the same transformer bl...

Michal Byra, Pawel Olszowiec, Grzegorz Stefanski et al. · 0 citations
#artificial intelligence Review May 2026

FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence

FraudBench is a multimodal benchmark for detecting AI-generated fraudulent refund evidence and shows that current MLLMs often recognize real-damaged evidence but fail on many fake-damaged subsets, with fake-damage detection rates far below the 50\% baseline on most generator subsets.

Xinyu Yan, Bo-Yang Chen, Jia-Ming Zhang et al. · 1 citation

LAGO: Language-Guided Adaptive Object-Region Focus for Zero-Shot Visual-Text Alignment

This work introduces LAGO (LAnguage-Guided adaptive Object-region focus), which reframes localized recognition as language-guided directed region discovery and addresses the circular dependence between recognizing a class and locating its supporting evidence, while preserving complementary local, contextual, and global...

Jun-Yi Hu, Qiji Zhou, Lei Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy

Vision-language-action (VLA) models can use visual prediction to anticipate future states, but dense visual features make the generative sequence grow with the number of camera views, prediction horizon, and encoder resolution. Whether such dense representations are necessary for effective control remains unclear. We i...

Zuojin Tang, Shengchao Yuan, Xiaoxin Bai et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning

Human visual reasoning is governed by active vision, a process where meta-cognitive control drives top-down goal-directed attention, dynamically routing foveal focus toward task-relevant details while maintaining peripheral awareness of the global scene. In contrast, modern Vision-Language Models (VLMs) process visual...

Brown Ebouky, Gabriele Carrino, Niccolo Avogaro et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Query-Dependent Use of Generated Descriptions for Reliable Visual Question Answering

Vision-Language Models (VLMs) hallucinate objects that are not present, and a growing line of work tries to curb this by feeding the model its own generated caption as auxiliary evidence -- assuming that a caption, once available, is something to consume. We show this fails: naively appending a caption can lower accura...

Zeshang Li, Shuoyang Zhang · 0 citations

DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams

The DRAGON dataset contains 11,664 annotated question instances from six diagram QA datasets, with a 2,445-instance test set carrying human-verified evidence annotations and a standardized evaluation framework, which supports future research on models that ground their predictions in visual evidence.

Anirudh Iyengar, Tampu Ravi Kumar, Gaurav Najpande et al. · 0 citations

SGP-SAM: Self-Gated Prompting for Transferring 3D Segment Anything Models to Lesion Segmentation

SGP-SAM, a self-gated prompting framework for efficient and effective transfer to 3D lesion segmentation, and a Zoom Loss that up-weights lesion-focused supervision by combining Dice and a voxel-balanced focal term to address small-lesion learning.

Ze-Quan Yao, Zi-Xuan Tang, Jie Ma et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

UniRect-CoT: Enhancing Generation in Unified Multimodal Models via Reflective Rectification with Inherent Understanding

Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable capability mismatch, where their understanding capability often outperforms their generation capability. This mismatch suggests that the model's rich internal knowledg...

Yibo Jiang, Tao Wu, Rui Jiang et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

CompDiff enables fair and zero shot medical image generation across demographic intersections through compositional diffusion

Medical image generators trained on imbalanced data can fail at demographic intersections absent from training. We introduce CompDiff, which encodes age, sex and race separately and composes supervised demographic tokens alongside clinical text. Across chest radiographs and fundus images, CompDiff improves overall and...

Mahmoud K. Ibrahim, Bart Elen, Chang Sun et al. · 0 citations
#artificial intelligence Preprint Mar 2026

Adaptive Weighted h-Transform Sampling for Coarse-Guided Visual Generation

A novel guided method is proposed by using the h-transform, a tool that can constrain stochastic processes (e.g., sampling process) under desired conditions and modify the transition probability at each sampling timestep by adding to the original differential equation with a drift function, which approximately steers t...

Yanghao Wang, Zi-Qi Jiang, Zhen Wang et al. · 3 citations · ⚡1

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.