Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Sep 2026

Reinforcement Learning from Intermediate Renders for Image-to-Code Generation

IR4RL is introduced, an RL framework with a token-level render-progress reward that turns changes between intermediate renders into localized feedback for the generated sequence, showing that intermediate rendering provides a simple and effective source of process supervision for RL post-training of image-to-code model...

Omri Kaduri, Kate Feingold, Phillip Isola et al. · 0 citations
#artificial intelligence Preprint Sep 2026

GenNVS: Geometry-enhanced Novel View Synthesis via Disentangled 3D Prior

Single-image novel view synthesis remains challenging because the underlying 3D geometry is highly ambiguous. Recent diffusion-based approaches produce plausible results, but they often struggle to preserve the geometric structure and spatial coherence of foreground objects. We present GenNVS, a framework for geometry-...

Ya-Jiao Xiong, You-Yu Luan, Xiao-Yu Zhou et al. · 0 citations
#artificial intelligence Review Sep 2026

ActionLens: Diagnosing Spatial-Temporal Binding Failures in Vision-Language Models

This work introduces ActionLens, a diagnostic benchmark of 6,701 multiple-choice video questions spanning five targeted diagnostics: transition detection, actor-specific identification, concurrent action binding, directed interaction reasoning, and gaze detection, and diagnostic measurements of these distinct failure m...

Gueter Josmy Faure, Min-Hung Chen, Hao-Ping Wang et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis

Vision-language (VL) pretraining using paired chest X-ray (CXR) images and radiology reports has shown strong potential for medical image understanding. However, existing methods often remain dependent on task-specific finetuning because radiology reports are lengthy, clinically dense, and difficult to align with simpl...

Hangyul Yoon, Hyungyung Lee, Edward Choi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Precise Editing and Flexible Referencing for Interactable Worlds

We present EditWorld, a video world model for precise editing and flexible referencing in interactable worlds. Existing video world models primarily focus on navigation, letting users explore generated worlds but offering limited control over how existing world content is modified. EditWorld extends world modeling from...

Xin-Yao Liao, Xian-Fang Zeng, Zhu Liang et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Modeling Whole-Slide Images as Dynamic Tumor Microenvironment Fields

Due to the gigapixel-scale nature of whole-slide images (WSIs), weakly supervised WSI analysis is commonly formulated as a multiple instance learning (MIL) problem, where patch-level features are aggregated into slide-level representations. However, diagnostic and prognostic evidence often arises from spatially coheren...

Lei Wu, Jiashuai Liu, Di Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Static to Dynamic: On-Policy Distillation from Image to Video Diffusion Models

MILD, a Motion-Preserving Image-to-Video Latent Distillation framework that transfers specialized image expertise while preserving pretrained video dynamics, is proposed and established as an effective route to improving video generation by drawing on the diverse and evolving capabilities of the image-generation ecosys...

Bing Jiang, Li Luo, Zi-Chao Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SyncRA: Learning Temporal Correspondence in Omni-Modal Models

Recent omni-modal models demonstrate strong perception of audio and visual inputs, yet often struggle to connect what they hear with what they see at the same moment. This weakness in temporal correspondence can cause models to associate spoken cues with the wrong visual scenes, producing plausible answers grounded in...

Ze-Long Xu, Yan Li, Wen-He Hu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SkillPE: Creativity-Oriented Cinematic Skill Evolution for Text-to-Video Prompt Engineering

Achieving high-quality, cinematic results in text-to-video generation remains challenging for non-experts, whose prompts often lack professional narrative and creative design. We propose SkillPE, a prompt engineering (PE) framework that evolves reusable cinematic skills from expert-authored seeds. SkillPE represents sh...

Yan-Wei Huang, Min Zhu, Shu-Jie Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?

Video-language models are increasingly used as judges of video understanding, both for evaluating model outputs and for training reward models. Whether their judgments remain reliable when the evidence is buried in day-long videos has yet to be established. Existing benchmarks cannot answer this. Their videos are typic...

Shayekh Bin Islam, Hwanjun Song · 0 citations
#artificial intelligence Review Sep 2026

See, Measure, and Reason: Learning Visually Grounded Reasoning in Pathology

Pathological assessment relies on recognizing fine-grained visual details in histological images. Vision-language models (VLMs) increasingly support pathology interpretation, yet their ability to perceive these details remains inadequate. This weakness leads to inaccurate cellular observations that can persist even whe...

Cheng-Yang Zhang, Wen-Chuan Zhang, Bo Li et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.