Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Sep 2026

Representation by Design in Generation: Cross-View Class-Token Alignment in Diffusion Transformers

Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas the models'own representations are often treated as a by-product of synthesis. We ask whether diffusion models can instead be trained to learn substantially stronger semanti...

Xiao-Yu Wu, Yi-Fei Wang, Chen Wei · 0 citations
#artificial intelligence Preprint Open access Sep 2026

PyroStack: A Multi-Band Spatio-Temporal Sub-Daily Dataset for Wildfires in the United States

Wildfires are an increasing hazard to ecosystems, air quality, and human systems, creating a growing need for datasets that support systematic development and evaluation of models for predicting fire spread across diverse landscapes. Effective prediction requires integrating meteorological conditions, fuels, vegetation...

Arya Kondur, Giosue Migliorini, Cameron Schmitt et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Think Before You Restore: Risk-Aware Manchu Manuscript Restoration with Stroke-Guided Attention

Full-page blind restoration of historical Manchu manuscripts is challenging due to scarce annotations, unknown degradation regions, and fragile connected strokes. Generic restoration models may improve visual quality but often modify intact content, leading to over-restoration. We propose SAGE-Restore (Stroke-Aware Gat...

Ming-Qiu Liang, Dongdong Wang, Si-Yang Lu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Mutually Adversarial Self-Training with Evolving Data for Unified Multimodal Models

MATE (Mutually Adversarial self-Training with Evolving data), a reinforcement-learning-based post-training framework in which the two branches instead challenge each other, and the challenges evolve as the model trains, turns the training into self-play in data space.

Wen-Tao Zhou, Wei-Jie Gan, Jia-Yun Wang · 0 citations
#artificial intelligence Review Sep 2026

PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute...

Vighnesh Subramaniam, B. Katz, Brian Cheung et al. · 0 citations
#artificial intelligence Preprint Sep 2026

One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models

Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing ga...

Aditya Sharma, Divya Saxena · 0 citations
#artificial intelligence Preprint Sep 2026

AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

A large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task, and a unified evaluation framework with a unified evaluation framework.

Tong-Tong Feng, Xin Wang, Hao-Ran Hou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method

Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowled...

Yun-Zhe Xu, Zhe Liu · 0 citations
#artificial intelligence Preprint Open access Sep 2026

PACT: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions

Single-token typed-decision models answer a schema question by reading the logits of a few one-letter answer codes at a single position: they are fast and return a probability for every allowed answer, but they are trained with plain cross-entropy that ignores most of the structure in their training data. We study such...

Yida Lin · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Brain-SAD: A Brain-Inspired Safe Autonomous Driving Control Framework with Dynamic Fear-Oriented Constraint on Dual-Policy

Constrained Reinforcement Learning has recently gained increasing attention in the field of Safe Autonomous Driving, where the general mechanism is to maximize the expected reward while keeping the overall action risk bounded. In this way, the safety issues arising in AD can be mitigated through constrained actions. Ho...

Huan Rong, Chao Yin, Anouar Imel et al. · 0 citations
#artificial intelligence Preprint Open access Sep 2026

Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution

Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However, execution traces contain only the evidence acquired by the current harness, leaving competing explanations for failure unresolved and limiting the basis for self-improv...

Bingjun Luo, Jialin Guo, Siqi Li · 0 citations
#artificial intelligence Preprint Sep 2026

Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics

Per-action evaluation and failure analysis highlight ambiguities from camera motion, delayed effects and imbalanced key-press frequencies that call for explicit modeling of 3D scene structure, long-term state and the adoption of proper losses in future implementations.

Abhishek Pillai, Ekta Prashnani, Joohwan Kim et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.