Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Sep 2026

From Modes to Memories: Characterizing the Scale-Space Dynamics of Diffusion Models

Diffusion models are typically viewed as stochastic processes that transform noise into data. We take a complementary perspective: a diffusion model defines a family of deterministic dynamical systems indexed by noise scale. At each fixed scale $\sigma$, we treat the denoiser as a self-map and study its dynamics. For a...

C. Amado, Marco Fumero, Francesco Locatello · 0 citations
#artificial intelligence Preprint Sep 2026

D-Scope: Decomposing and Steering Diffusion Transformers with Sparse Autoencoders

Sparse autoencoders (SAEs) reveal visual structure in diffusion transformers (DiTs), but interpreting a feature does not establish whether it can be used to control generation. We introduce D-Scope (Diffusion Scope), a framework that connects feature interpretation to generation control through shared visual evidence....

Xinyue Xu, Jiahao Zhang, Li-Jie Hu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generati...

Qi-Ze Yu, Lian-Rui Fan, Bo-Yu Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply...

Qize Yu, Lianrui Fan, Bowen Ping et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Steering Fields: Adaptive Vector Fields for Safe Image Generation and Beyond

As state-of-the-art text-to-image flow models achieve near-photorealistic quality, controlling their outputs, e.g., suppressing harmful content while promoting benign alternatives, has become a central challenge. The current steering paradigm consists of adding a global steering vector to selected activations. While fu...

Simone Facchiano, Jan Eric Lenssen, Bernt Schiele et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models

Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack mechanisms. In this...

Jun-Jian Li, Xiao-Long Liu, Peng Sun et al. · 1 citation
#artificial intelligence Preprint Sep 2026

PartiCam: Camera Controlled Video Generation with Reward Guidance

We present PartiCam, a training-free Particle filtering rooted method for improved Camera controlled video generation. Generating videos that follow a precisely specified camera trajectory remains challenging for large video diffusion models. Training-free approaches are backbone-agnostic and avoid the need to construc...

Amine Ouasfi, Runjia Li, Junlin Han et al. · 0 citations
#artificial intelligence Preprint Sep 2026

OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning

Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluated and insufficiently elicited. We addre...

Junming Lin, Yuxuan Wang, Zhen-Xin Lei et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models

Online reinforcement learning has been extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where exploration noise is...

Yu Shu, Chao-Chao Lu · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification

Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a multi-task Deep Learning framework for MRI-based glioma diagnosis that performs tumor segmentation and predicts IDH mutation status, 1p/19q co-deletion status, and tumor grade...

Gonzalo Esteban Mosquera Rojas, Sebastian R. van der Voort, Carolin M. Pirkl et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Rethinking Multi-Image Re-Representation in Multi-Image Understanding

Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them. We study this problem through multi-image re-representation, viewing prompted Chain-of-Thought reasoning and agentic visual tool use as different ways of re-orga...

Gengyuan Zhang, Xiao Han, Xinyu Xie et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Emergent Multi-View Geometry Through Self-Distillation

Over a century ago, Henri Poincar\'e argued that a motionless observer cannot acquire the notion of space. Yet, most visual representation learning methods operate on individual images, while those that leverage multiple views rely on RGB reconstruction, entangling geometry with appearance. We propose Poincar3, a self-...

David Nordström, T. Loiseau, Vincent Lepetit et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.