Skip to content

Category

computer vision

3,022 papers

#artificial intelligence Preprint Open access Oct 2026

Language-Conditioned World Modeling for Visual Navigation

Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, language-conditioned visual navigation (LCVN), in which an embodied agent must follow a natural language instruction given only an initial egocentric observation. Without access to goal...

Yifei Dong, Fengyi Wu, Yilong Dai et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Delta-K: Boosting Multi-Instance Generation via Cross-Attention Augmentation

While Diffusion Models excel in text-to-image synthesis, they frequently suffer from catastrophic concept omission when generating complex multi-instance scenes. Existing training-free methods attempt to resolve this by rescaling attention maps, which merely exacerbates unstructured noise without establishing coherent...

Zitong Wang, Zijun Shen, Haohao Xu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

ITO: Multi-View Alignment and Training-Time Fusion for Image-Text Pretraining

Image--text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality rather than by semantics. We propose ITO, a framework addressing this limitation through two complementary mechanisms with d...

Hanpeng Liu, Zidan Wang, Shuoxi Zhang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

PISCO: Precise Video Instance Insertion with Sparse Control

The landscape of AI video generation is undergoing a pivotal shift: moving beyond general generation - which relies on exhaustive prompt-engineering and "cherry-picking" - towards fine-grained, controllable generation and high-fidelity post-processing. In professional AI-assisted filmmaking, it is crucial to perform pr...

Xiangbo Gao, Renjie Li, Xinghao Chen et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

Edge-Aware and Content-Adaptive Infrared Gas Leak Detection for Industrial Safety Monitoring

Infrared gas leak detection is important for industrial safety and environmental monitoring, but automatic detection remains challenging because gas plumes are often faint, small, semi-transparent, and weakly bounded. This study proposes an Edge-Aware and Content-Adaptive Feature Fusion Detector (ECAF-Det) for infrared...

Dongsheng Li, Tianli Ma, Siling Wang et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

From Compound Figures to Medical Multi-image Reasoning: Scaling Multimodal Large Language Models with Biomedical Literature

Multimodal large language models (MLLMs) are increasingly capable in medical imaging, yet most focus on single-image settings. Clinical interpretation often requires integrating evidence across multiple images, such as different modalities, views, or time points. However, large-scale medical multi-image data and traini...

Zhen Chen, Yihang Fu, Rong Zhou et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

FoCLIP: A Feature-Space Misalignment Framework for CLIP-Based Image Manipulation and Detection

The well-aligned attribute of CLIP-based models enables its effective application like CLIPscore as a widely adopted image quality assessment metric. However, such a CLIP-based metric is vulnerable for its delicate multimodal alignment. In this work, we propose FoCLIP, a feature-space misalignment framework for fooling...

Yulin Chen, Zeyuan Wang, Tianyuan Yu et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

3D Software Synthesis Driven by Constraint-Expressive Intermediate Representation

Graphical user interface (UI) software has undergone a fundamental transformation from traditional two-dimensional (2D) desktop/web/mobile interfaces to spatial three-dimensional (3D) environments. While existing work has made remarkable success in automated 2D software generation, such as HTML/CSS and mobile app inter...

Shuqing Li, Anson Y. Lam, Yun Peng et al. · 0 citations
#artificial intelligence Preprint Open access Oct 2026

AneumoBench: A Source-Linked Benchmark for Synthetic-Geometry Transfer in Aneurysm CFD

Scientific machine learning uses simulation data to train surrogate models for fast physical-field prediction across geometries. Local shape editing can expand limited geometry collections, but whether its variants improve prediction on unseen geometries, and how to allocate them across sources, require controlled eval...

Xigui Li, Yuanye Zhou, Feiyang Xiao et al. · 0 citations
#artificial intelligence Review Sep 2026

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

Recent video generation is increasingly realistic and controllable, yet video editing remains less developed, particularly for precise local edits that must preserve the original scene dynamics. Video scene text editing replaces text on scene surfaces, such as storefront signs, whiteboards, and product labels, while pr...

Xing-Hao Chen, Xiang-Bo Gao, Jiong-Ze Yu et al. · 0 citations
#artificial intelligence Review Sep 2026

MatLoom: Layered Text-to-Material Generation in a Compact Program Space

Material generation should produce not only an appearance, but also the rules that construct it. We introduce MatLoom, a compact, layer-oriented language for text-to-material generation with pretrained language models. Each program composes alpha-masked layers whose shared spatial expressions define coverage and physic...

Anson Y. Lam, Shu-Qing Li, M. Lyu · 0 citations
#artificial intelligence Preprint Sep 2026

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privil...

Yong Du, Tong-I Chen, Zhengxi Lu et al. · 0 citations

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.