Skip to content

Category

computer vision

3,022 papers

#machine learning Preprint Oct 2026

Scale-Recursive Rectified Flows for Few-Step Precipitation Ensembles

Fine-resolution precipitation estimates support flood risk assessment and water management, but coarse satellite products cannot resolve rainfall within each grid cell. Generative models address this ambiguity by producing ensembles of plausible high-resolution rainfall fields. Among these models, rectified flows gener...

Shunya Nagashima, Takumi Bannai · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text-3D interaction remains largely implicit. Existing methods concatenate text and 3D tokens into a flat sequence and rely on self-attention, collapsing coarse structural cues and fine geomet...

Tianjiao Yu, Xinzhuo Li, Yifan Shen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

GUI Agents for Continual Game Generation

Generating a game is not the same as making one playable. Existing code-generation approaches often translate a prompt directly into an artifact, leaving interaction-level failures undetected. We argue that game generation requires a player and study two roles for graphical user interface (GUI) agents. First, we introd...

Yixu Huang, Bo Li, Na Li et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

DuetMoE: Coupling Inter- and Intra-Subgroup Robustness for Fair Medical Image Analysis

As medical AI expands across diverse healthcare settings worldwide, equitable performance across patient populations is becoming essential to trustworthy clinical use. Fairness in medical image analysis is often evaluated through average performance across predefined subgroups, yet similar subgroup averages can conceal...

Yiqi Tian, Jinwoong Park, Sangjoon Park et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity

Automated fact-checking is a crucial task that supports a responsible information ecosystem. While recent research has progressed from text-only to multimodal fact-checking, a prevailing assumption is that incorporating visual evidence universally improves verification accuracy. In this work, we challenge this assumpti...

Jaeyoon Jung, Yejun Yoon, Kunwoo Park · 0 citations
#artificial intelligence Preprint Open access Oct 2026

The Effective Depth Paradox: Topology and Trainability in Deep CNNs

This paper presents a controlled comparative study of convolutional neural network (CNN) topology and image classification performance across the architectural families VGG, ResNet, and GoogLeNet, evaluated on CIFAR-10 under a unified training protocol. We formalize the distinction between nominal depth ($D_{\mathrm{no...

Manfred M. Fischer, Joshua Pitts · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Extended to Reality: Prompt Injection in 3D Environments

Multimodal large language models (MLLMs) have advanced the capabilities to interpret and act on visual input in 3D environments, empowering diverse applications such as robotics and situated conversational agents. When MLLMs reason over camera-captured views of the physical world, a new attack surface emerges: an attac...

Zhuoheng Li, Ying Chen · 0 citations
#artificial intelligence Preprint Open access Oct 2026

VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation

Autoregressive (AR) models have recently shown strong performance in image generation, where a critical component is the visual tokenizer (VT) that maps continuous pixel inputs to discrete token sequences. The quality of the VT largely defines the upper bound of AR model performance. However, current discrete VTs fall...

Huawei Lin, Tony Geng, Zhaozhuo Xu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

VIDiff: Translating Videos via Multi-Modal Instructions with Diffusion Models

Diffusion models have achieved significant success in image and video generation. This motivates a growing interest in video editing tasks, where videos are edited according to provided text descriptions. However, most existing approaches only focus on video editing for short clips and rely on time-consuming tuning or...

Zhen Xing, Shuyuan Tu, Qi Dai et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis

This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack...

Keerthi Kaashyap, D. Anthony, Akshay Krishnan et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstra...

Ruihong Shen, \v{Z}iga Kova\v{c}i\v{c}, Peter Kulits et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

What Should World Models Forget? Stratified Retention for Continual Adaptation

Continual learning treats degradation on previously seen data as evidence of failure, a convention inherited from settings with a stationary prediction target, where a correct label remains correct indefinitely. World models do not satisfy this condition. Their prediction target is the environment, which changes, so kn...

Nishit Anand, Ramani Duraiswami, Dinesh Manocha · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.