Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

Visual Jev Rewards: Reference-Bound Verification for Multi-Subject Image Generation

Multi-subject image generation requires rewards that verify whether requested attributes, actions, and relations hold for the specified reference subjects. Subject presence alone does not establish that the correct subjects participate in a requested interaction. We present reference-bound Visual Jev rewards that turn...

Baoteng Li, Wenzhuo Wu, Kongming Liang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Kuration SDK: Addressing the Virtual2Real Gap via Data Curation

Benchmarks for measuring the quality of action-conditioned world models are still evolving and shifting away from visual similarity-based metrics to action-semantic and physically-grounded metrics. However, for domain and task-agnostic action-conditioned world model training, existing benchmarks provide a limited signa...

Nirmit Desai, Eric Song, Mayank Sengupta et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

LeCuration: A Tiny World Model as a Data Curation Multi-Tool

Many applications of physical AI run within finite or closed physical worlds with a limited set of physical laws governing object behavior. Examples include robots working in a warehouse and agents moving around in a video game. In order to better organize, filter, and curate data for physical AI applications, we propo...

Mayank Sengupta, Nirmit Desai, Eric Song et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Consistent Distribution Matching for Data-Free Diffusion Distillation

Flow and diffusion models suffer from slow inference due to computationally expensive numerical integration. Distillation provides a promising way for a student model to learn from a teacher's dynamics, enabling one-step or few-step generation. However, existing methods often depend on curated distillation datasets, co...

Yuxiang Fu, Qi Yan, Zike Wu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

A Deterministic Evidence Layer for Vision-Language Autism Screening from Naturalistic Home Video

Autism spectrum disorder (ASD) is diagnosed through specialist observation of a child's social behavior, and access to that expertise is the bottleneck for early identification. Vision-language models (VLMs) describe a child's behavior from video well; the verdict drawn from the description is unstable: at temperature~...

Wenqi Li, Mindi Ruan, Chuanbo Hu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Do Vision Models Learn Physical Constraints or Rendering Shortcuts? A Counterfactual Benchmark for Grounded Physical Consistency

Modern image editing models can satisfy a text instruction while breaking the physics of the edited scene. A new object may cast no shadow, a mirror may fail to reflect visible geometry, or an object may float above a surface that should support it. We study physical plausibility diagnosis, detecting whether an edited...

M. Moein Esfahani, Sepehr Salem, Mohammed Alser et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Visual Memory Attacks Can Persist Through The KV Cache

Modern language model systems operate autonomously over increasingly long contexts containing untrusted text and images. Can an adversarial input continue to steer a model even after that input is removed from its context? We show that attacks can be trained to persist through the key/value (KV) cache of subsequent tok...

David Dobre, Leo Schwinn, Gauthier Gidel et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Personalize at Test Time: Learning User Preferences for Image Generation

Diffusion models can generate high-quality images, yet aligning their outputs with individual user preferences remains challenging. A key bottleneck is accurately modeling diverse user preferences from limited feedback. Existing approaches often rely on labor-intensive manual preference annotations or vision-language m...

Jiamu Bai, Jiaming Hu, Yanhong Wu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

RACER: Reflective Agent Coupling Query Interpretation and Tool-Based Retrieval for Frame Selection in Long Video Understanding

Video large language models (Vid-LLMs) excel at diverse video-language tasks by reasoning over selected frames. However, frame selection for long videos remains challenging, as it requires retrieving relevant frames distributed across segments from a large candidate pool given complex queries. This paper investigates d...

Yiyang Huang, Yitian Zhang, Yizhou Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

One-Slide Calibration of Pathology Foundation Models

Scanner variation changes how pathology foundation models represent the same tissue. We introduce SlideRuler, which uses regions within a slide as internal controls to estimate and correct acquisition-induced shifts in other regions. A transfer map learned from paired rescans enables calibration from a single scan at i...

Ming Ren Hou, Tianyi Huang · 0 citations
#artificial intelligence Preprint Open access Oct 2026

STRIDE: Spatial-Temporal Representation for Interval-conditioned Disease Evolution in Longitudinal Glioblastoma MRI

Glioblastoma (GBM), an aggressive primary brain tumor, is routinely monitored with longitudinal MRI after treatment. Distinguishing stable disease (SD), pseudoprogression (PsP), and true progression (TP) remains challenging because these states can show overlapping MRI appearances despite different temporal trajectorie...

Wenhao Guo, Changchang Yin, Pierre Giglio et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Deep Learning for Longitudinal Medical Imaging: A Scoping Review

Longitudinal medical imaging analysis is a cornerstone of modern medical practice and patient care. Deep learning applied to longitudinal imaging offers wide potential to enhance diagnosis and track disease progression by capturing spatial changes over time. With major advances in single-timepoint deep learning for ima...

Francesca Mussa, Divyanshu Tak, Atlas H. Avval et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.